Slide show direction switching method and device
By using a large language model to analyze and group semantic themes in slide conversion, combined with dynamic layout adjustment technology, the problems of low automation, poor content adaptability and insufficient intelligence in the existing technology are solved, and efficient and intelligent slide horizontal to vertical conversion is achieved.
Patent Information
- Application Number
- CN202510363504.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-03-26
AI Technical Summary
In the process of horizontal to vertical conversion of slides, the prior art has problems such as low degree of automation, poor content adaptability and insufficient intelligence, which cannot meet the needs of multi-scene slide display.
By introducing a large language model to analyze and intelligently group page elements, combined with dynamic layout adjustment technology, a more efficient display layout is achieved. The specific steps include extracting page display information, determining text input information for the large language model, grouping semantic topics, and adjusting page size and element arrangement.
It realizes the semantic relationship of page elements automatically and intelligently adjusts the page layout to ensure that the converted page content is uniformly distributed, the logic is clear and visually beautiful, and improves the conversion efficiency and display effect.
Smart Images

Figure CN119884397B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer software technology, and in particular to a method and device for switching the direction of slide presentation. Background Art
[0002] When presenting slides, horizontal to vertical conversion is a common requirement, especially when presenting on mobile devices and small-screen devices, vertical layout can better adapt to users' reading habits. However, existing slide horizontal to vertical conversion technologies, such as manual adjustment, simple zooming, and preset template conversion, have many limitations.
[0003] Although the manual adjustment method can accurately modify the page layout according to user needs, its operation is cumbersome and time-consuming. Especially when the page elements are complex, the user needs to adjust the position and size of each element one by one, which easily leads to inconsistent layout and affects the overall aesthetics and logical consistency. In addition, the manual adjustment of different pages may lack a unified standard, further exacerbating the problem of inconsistent layout style.
[0004] Simple zooming is to scale the content of the horizontal slide in proportion. Although it is simple to implement and can quickly complete the conversion, this method cannot distinguish the semantic weight of different content, resulting in text being too small and images being blurred. Especially when displayed on small-screen devices, users need to zoom in frequently to view, which seriously affects the reading experience and information transmission effect.
[0005] The purpose of preset template conversion is to reorganize horizontal slide content into vertical format through predefined templates. Although it simplifies the conversion process, the template types are limited and it is difficult to adapt to the needs of complex content scenarios. When the slide content is complex and the page structure is diverse, the template cannot be flexibly adjusted, resulting in conversion results that do not meet expectations. A lot of manual adjustments are still required, which reduces the degree of automation.
[0006] In addition, the existing technology lacks intelligent semantic analysis and cannot identify the semantic relationship between page elements and their visual priority. For example, the relative positions of different contents such as titles, texts and auxiliary elements cannot be intelligently adjusted, resulting in a lack of logic and visual hierarchy in the content layout, affecting the overall display effect.
[0007] In summary, the existing technology has deficiencies in terms of automation, content adaptability and intelligence level, and cannot meet the needs of multi-scenario slide show presentations.
[0008] Therefore, how to achieve adaptive layout of content during the transition from horizontal to vertical orientation of slides, improve display effects and increase conversion efficiency has become a technical problem that needs to be solved urgently. Summary of the invention
[0009] The embodiments of the present application provide a method and device for switching the display direction of slides, aiming to solve the technical problem of poor horizontal to vertical conversion effect of slides in related technologies. By using a large language model to perform semantic theme analysis and intelligent grouping of page elements, content with the same semantic theme is automatically grouped together, thereby achieving a more efficient display layout during the horizontal to vertical conversion of slides.
[0010] In a first aspect, an embodiment of the present application provides a method for switching a slide presentation direction, comprising:
[0011] Extracting page display information from the target slide, wherein the page display information is used to reflect the display characteristics of the current shape page and the page elements;
[0012] Determining text input information of a large language model based on the page display information;
[0013] Inputting the text input information and a preset prompt for the text input information into the large language model, so that the large language model outputs page element classification information for grouping page elements according to semantic themes, wherein page elements with the same semantic theme are in the same group;
[0014] Determining the size of the target shape page based on the page display information and a predetermined page size adjustment rule;
[0015] Arrangement information of the page elements in the current shape page within the target shape page is determined based on the size of the target shape page, the page element classification information and a predetermined page element arrangement rule.
[0016] In one embodiment of the present application, optionally, the page display information includes the page display characteristics of the current shape page and the element display characteristics of the page elements in the current shape page, wherein
[0017] The page display characteristics include: one or more of resolution, size, page elements and page annotations;
[0018] The element display characteristics of the page element include: one or more of: position, level, resource path, size, content, type and element description text.
[0019] In one embodiment of the present application, optionally, extracting the display information of the current shape page in the target slide includes:
[0020] Locate the OfficeOpenXML file of the target slide <p:sldsz>Label;
[0021] Get the <p:sldsz>The cx attribute and the cy attribute in the tag, wherein the cx attribute and the cy attribute represent the width and height of the current shape page of the target slide respectively;
[0022] Determining the size and resolution of the current shape page based on the cx attribute and the cy attribute; and
[0023] Locate the OfficeOpenXML file of the target slide <p:csld>Label and Description <p:csld>Subtags of a tag;
[0024] Based on the <p:csld>Label and Description <p:csld>The sub-tag of the tag determines the page elements in the current shape page and the element types of the page elements, wherein the types include: one or more of text box, image, graphic, and background;
[0025] Locate the OfficeOpenXML file of the target slide <pp:notes>Label;
[0026] The <pp:txbody>Label and Description <pp:txbody>The text information stored in each sub-tag of the tag is determined as the page annotation of the current shape page;
[0027] Locate the OfficeOpenXML file of the target slide <a:t>Label;
[0028] Based on the element type of each of the page elements, <a:t>The element display feature corresponding to each of the page elements is extracted from the text content stored in the tag.
[0029] In one embodiment of the present application, optionally, it further includes:
[0030] For the image in the page element, generating semantic description information of the image based on the image description model;
[0031] Then, determining the text input information of the large language model based on the page display information includes:
[0032] The semantic description information of the image is added to the text input information of the large language model.
[0033] In one embodiment of the present application, optionally, before determining the text input information of the large language model based on the page display information, the method further includes:
[0034] Convert the page display information into a JSON format file;
[0035] Then, determining the text input information of the large language model based on the page display information includes:
[0036] Extracting text information, image description and context description associated with page elements from the JSON format file obtained by converting the page display information as text input information of the large language model;
[0037] After determining the arrangement information of the page elements in the current shape page in the target shape page, the method further includes:
[0038] The arrangement information is converted into a JSON format file.
[0039] In one embodiment of the present application, optionally, the page element classification information output by the large language model includes: a label of the page element and a function type in the target slide;
[0040] After the large language model is made to output page element classification information for grouping page elements according to semantic topics, the method further includes:
[0041] The page element classification information is standardized, wherein the standardization includes deleting redundant tags, redundant function types and formatting symbols.
[0042] In one embodiment of the present application, optionally, determining the size of the target shape page based on the page display information and a predetermined page size adjustment rule includes:
[0043] Determining a first ratio of a width of the current shape page to a width of the target shape page, and determining a second ratio of a height of the current shape page to a height of the target shape page;
[0044] Determine the minimum value of the first ratio and the second ratio as the lower limit value of the scaling ratio for converting the current shape page into the target shape page, and set the upper limit value of the scaling ratio to 1;
[0045] In the zoom ratio range from the zoom lower limit value to the zoom upper limit value, searching for the maximum feasible zoom ratio that satisfies the preset layout rule by binary search, wherein the preset layout rule indicates that all page elements in the current shape page can be successfully displayed in the target shape page without changing their size;
[0046] Based on the size of the current shape page and the maximum feasible zoom ratio, the size of the target shape page is determined.
[0047] In one embodiment of the present application, optionally, determining arrangement information of page elements in the current shape page within the target shape page based on the size of the target shape page, the page element classification information and a predetermined page element arrangement rule includes:
[0048] Initializing page layout parameters of the target shape page, wherein the page layout parameters include page margins, row spacing, column spacing, and element arrangement direction;
[0049] Traverse all the page elements according to the element arrangement direction, and during the traversal process,
[0050] Obtaining the priorities of all the page elements, and allocating page space to page elements with higher priorities first when arranging the elements;
[0051] If the accumulated width of the elements in the current row exceeds the width limit of the page margin, a line break operation is triggered;
[0052] When performing a line break operation, the remaining arrangeable height of the target shape page is determined based on the page margin and the cumulative height of the elements. If the remaining arrangeable height is less than the height of the page elements to be arranged, a paging operation is triggered, or a trimming operation is performed on the page elements to be arranged.
[0053] In one embodiment of the present application, optionally, it further includes:
[0054] If the background element in the page element is a texture background or a gradient background, the background element is synchronously scaled according to the maximum feasible scaling ratio;
[0055] If the background element in the page element is not a texture background or a gradient background, then, based on the respective sizes of the current shape page and the target shape page and the position of the background element, the individual scaling ratio of the background element is determined, wherein:
[0056] The position of the background element includes the coordinates of the upper left corner vertex and the lower right corner vertex of the background element;
[0057] The individual scaling ratio of the background element is consistent with the height scaling ratio or the width scaling ratio of the current shape page;
[0058] Determine the coordinates of the lower right corner vertex of the background element in the target shape page based on the coordinates of the lower right corner vertex of the background element in the current shape page;
[0059] Determine the upper left corner vertex coordinates of the background element in the target shape page according to the individual scaling ratio of the background element and the lower right corner vertex coordinates of the background element in the target shape page;
[0060] If the upper left corner vertex coordinates extend outside the target shape page, a cropping operation is performed on the background element so that the upper left corner vertex coordinates of the cropped background element are located within the target shape page, wherein the cropped background element is the image center area of the background element before cropping.
[0061] In a second aspect, an embodiment of the present application provides a slide display direction switching device, comprising: a page display information extraction unit, used to extract page display information in a target slide, wherein the page display information is used to reflect the display characteristics presented by the current shape page and the page elements respectively; a text input information determination unit, used to determine the text input information of a large language model based on the page display information; an LLM processing unit, used to input the text input information and a preset prompt for the text input information into the large language model, so that the large language model outputs page element classification information for grouping page elements according to semantic themes, wherein page elements with the same semantic theme are in the same group; a size adjustment unit, used to determine the size of a target shape page based on the page display information and a predetermined page size adjustment rule; an element rearrangement unit, used to determine the arrangement information of the page elements in the current shape page within the target shape page based on the size of the target shape page, the page element classification information and a predetermined page element arrangement rule.
[0062] In one embodiment of the present application, optionally, the page display information includes page display features of the current shape page and element display features of page elements within the current shape page, wherein the page display features include: one or more of resolution, size, page elements and page annotations; the element display features of the page elements include: one or more of position, hierarchy, resource path, size, content, type and element description text.
[0063] In one embodiment of the present application, optionally, the page display information extraction unit includes: a first extraction unit for locating the OfficeOpenXML file of the target slide <p:sldsz>Tags; Get the <p:sldsz>The cx attribute and the cy attribute in the tag, wherein the cx attribute and the cy attribute represent the width and height of the current shape page of the target slide respectively; based on the cx attribute and the cy attribute, the size and resolution of the current shape page are determined; the second extraction unit is used to locate the OfficeOpenXML file of the target slide <p:csld>Label and Description <p:csld>Subtags of the tag; based on the <p:csld>Label and Description <p:csld>The subtag of the tag determines the page elements in the current shape page and the element types of the page elements, wherein the types include: one or more of text boxes, images, graphics, and backgrounds; the third extraction unit is used to locate the OfficeOpenXML file of the target slide <pp:notes>label; <pp:txbody>Label and Description <pp:txbody>The text information stored in each of the sub-tags of the tag is determined as the page annotation of the current shape page; the fourth extraction unit is used to locate the OfficeOpenXML file of the target slide <a:t>Tag; based on the element type of each of the page elements, in the <a:t>The element display feature corresponding to each of the page elements is extracted from the text content stored in the tag.
[0064] In one embodiment of the present application, optionally, the device also includes: an image description unit, which is used to generate semantic description information of the image in the page element based on an image description model; then the text input information determination unit is also used to: add the semantic description information of the image to the text input information of the large language model.
[0065] In one embodiment of the present application, optionally, the device also includes: a first format conversion unit, used to convert the page display information into a JSON format file before the text input information determination unit determines the text input information of the large language model; a second format conversion unit, used to convert the arrangement information into a JSON format file after determining the arrangement information of the page elements in the current shape page in the target shape page; the text input information determination unit is used to: extract text information, image description and context description associated with the page elements in the JSON format file obtained by converting the page display information as the text input information of the large language model.
[0066] In one embodiment of the present application, optionally, the page element classification information output by the large language model includes: the label of the page element and the functional type in the target slide; the device also includes: a standardization processing unit, which is used to standardize the page element classification information after the LLM processing unit causes the large language model to output the page element classification information that groups the page elements according to semantic themes, wherein the standardization processing includes deleting redundant labels, redundant functional types and formatting symbols.
[0067] In one embodiment of the present application, optionally, the size adjustment unit includes: determining a first ratio of the width of the current shape page to the width of the target shape page, and determining a second ratio of the height of the current shape page to the height of the target shape page; determining the minimum value of the first ratio and the second ratio as the lower limit value of the scaling ratio for converting the current shape page to the target shape page, and setting the upper limit value of the scaling ratio to 1; within the scaling ratio range from the lower limit value to the upper limit value, searching for the maximum feasible scaling ratio that satisfies the preset layout rules by binary search, wherein the preset layout rules indicate that the page elements in the current shape page can be smoothly displayed in the target shape page without changing the size; determining the size of the target shape page based on the size of the current shape page and the maximum feasible scaling ratio.
[0068] In one embodiment of the present application, optionally, the element rearrangement unit is used to: initialize the page layout parameters of the target shape page, wherein the page layout parameters include page margins, row spacing, column spacing, and element arrangement direction; traverse all the page elements according to the element arrangement direction, and during the traversal process, obtain the priorities of all the page elements, and preferentially allocate page space to page elements with higher priorities when arranging elements; if the cumulative width of the elements in the current row exceeds the width limit of the page margins, trigger a line break operation; when performing a line break operation, determine the remaining arrangeable height of the target shape page based on the page margins and the cumulative height of the elements, and if the remaining arrangeable height is less than the height of the page elements to be arranged, trigger a paging operation, or perform a trimming operation on the page elements to be arranged.
[0069] In one embodiment of the present application, optionally, the device further includes:
[0070] The background adjustment unit is used to synchronously scale the background element according to the maximum feasible scaling ratio if the background element in the page element is a texture background or a gradient background; if the background element in the page element is not a texture background or a gradient background, determine the individual scaling ratio of the background element based on the respective sizes of the current shape page and the target shape page and the position of the background element, wherein the position of the background element includes the coordinates of the upper left corner vertex and the lower right corner vertex of the background element; the individual scaling ratio of the background element is consistent with the height scaling ratio or the width scaling ratio of the current shape page; based The lower right corner vertex coordinates of the background element in the current shape page are used to determine the lower right corner vertex coordinates of the background element in the target shape page; the upper left corner vertex coordinates of the background element in the target shape page are determined according to the individual scaling ratio of the background element and the lower right corner vertex coordinates of the background element in the target shape page; if the upper left corner vertex coordinates extend outside the target shape page, a cropping operation is performed on the background element so that the upper left corner vertex coordinates of the cropped background element are located within the target shape page, wherein the cropped background element is the image center area of the background element before cropping.
[0071] In a third aspect, an embodiment of the present application provides a computer device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method described in the first aspect above.
[0072] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the method described in the first aspect above.
[0073] The above technical solution aims at the technical problem of poor horizontal to vertical conversion effect of slides in related technologies, and solves the problems of low manual adjustment efficiency, reduced readability due to simple scaling, insufficient flexibility of preset templates, and lack of intelligent semantic analysis in the existing technology. By introducing a large language model and dynamic layout adjustment technology, the goal of automatically identifying the semantic relationship of page elements and intelligently adjusting the page layout can be achieved, ensuring that the content of the converted page is evenly distributed, logically clear, and visually beautiful, achieving automatic and efficient adaptive layout, high conversion efficiency, and high-quality conversion effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0075] Figure 1 A flow chart showing a method for switching slide presentation direction according to an embodiment of the present application is shown;
[0076] Figure 2 A flow chart showing a method for switching slide presentation direction according to another embodiment of the present application is shown;
[0077] Figure 3 A schematic diagram showing a slide display direction switching device according to an embodiment of the present application is shown;
[0078] Figure 4 A block diagram of a computer device according to an embodiment of the present application is shown;
[0079] Figure 5 A block diagram of a computer device according to another embodiment of the present application is shown;
[0080] Figure 6 A block diagram of a computer device according to yet another embodiment of the present application is shown. DETAILED DESCRIPTION
[0081] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0082] Figure 1 A flow chart of a method for switching slide presentation direction according to an embodiment of the present application is shown.
[0083] like Figure 1 As shown, a method for switching slide presentation direction according to an embodiment of the present application includes:
[0084] Step 102: extract page display information from the target slide.
[0085] The page display information is used to reflect the display features of the current shape page and the page elements, including the page display features of the current shape page and the element display features of the page elements in the current shape page. Specifically, the page display features include: one or more of resolution, size, page elements and page annotations; the element display features of the page elements include: one or more of position, hierarchy, resource path, size, content, type and element description text.
[0086] It should be noted that, in this context, the current shape page refers to the original slide page to be converted, and the target shape page refers to the converted slide page.
[0087] In a possible design, the current shape page is in a horizontal layout, and the target shape page is in a vertical layout.
[0088] In another possible design, the current shape page is in a vertical layout, and the target shape page is in a horizontal layout.
[0089] By extracting page display information, the layout structure and content features of the target slide can be fully understood, providing basic data support for subsequent semantic analysis and layout adjustment. Specifically, the extraction of page display information enables the system to identify the overall structure of the page and the detailed information of each element, thereby providing accurate input data for subsequent intelligent adjustments. This step ensures that subsequent processing can be based on accurate page information, avoiding layout deviations caused by incomplete or erroneous information.
[0090] In a possible design, the OfficeOpenXML file of the target slide can be located. <p:sldsz>Tags; Get the <p:sldsz>The cx attribute and the cy attribute in the tag, wherein the cx attribute and the cy attribute respectively represent the width and height of the current shape page of the target slide; based on the cx attribute and the cy attribute, the size and resolution of the current shape page are determined.
[0091] OfficeOpenXML is an XML-based file format used to represent electronic documents and is widely used in office software suites. The present invention extracts slide information by parsing OfficeOpenXML files. For relevant specifications, see ISO / IEC 29500 standard.
[0092] This step provides basic data support for subsequent page layout adjustment and scaling. By accurately obtaining the page size, the scaling ratio and layout parameters can be dynamically calculated according to the size requirements of the target page, ensuring that the converted page content can adapt to the new display environment, avoiding the layout disorder caused by size mismatch in traditional methods.
[0093] In another possible design, the OfficeOpenXML file of the target slide can be located. <p:csld>Label and Description <p:csld>Subtags of the tag; based on the <p:csld>Label and Description <p:csld>The sub-tag of the tag determines the page elements in the current shape page and the element type of the page elements, wherein the type includes: one or more of: text box, image, graphic, background.
[0094] By parsing <p:csld>Tags and their sub-tags, the system can identify all page elements and their types (such as text boxes, images, graphics, backgrounds, etc.) contained in the current page. This step provides detailed element information for subsequent semantic analysis and layout optimization. By accurately identifying the types of page elements, differentiated processing strategies can be adopted according to different types of elements, such as text typesetting optimization for text boxes, scaling and cropping for images, and level adjustment for graphics. This refined processing method improves the accuracy of page conversion and display effects.
[0095] In another possible design, the OfficeOpenXML file of the target slide can be located. <pp:notes>Label; the <pp:txbody>Label and Description <pp:txbody>The text information stored in each of the sub-tags of the tag is determined as the page annotation of the current shape page.
[0096] That is, by analyzing <pp:notes>The tag and its subtags can extract page annotation information. Page annotations usually contain supplementary explanations or contextual information about the slide content, which can provide additional contextual support for subsequent semantic analysis. By utilizing page annotations, the system can more accurately understand the semantic relationship of page elements and generate a more reasonable layout adjustment plan. This step enhances the intelligence level of the system and improves the logic and readability of the conversion results.
[0097] In another possible design, the OfficeOpenXML file of the target slide can also be located. <a:t>Tag; based on the element type of each of the page elements, in the <a:t>The element display feature corresponding to each of the page elements is extracted from the text content stored in the tag.
[0098] By parsing <a:t>Tags can extract the display features of each page element, including position, size, content, description text, etc. This step provides detailed element attribute information for subsequent layout optimization. By accurately obtaining the display features of each element, it is possible to dynamically adjust its arrangement in the target page according to the element's size and position information to ensure that the content is evenly distributed and the logic is clear. For example, for text boxes, the text layout can be adjusted according to its content and size; for pictures, they can be scaled and cropped according to their size and resource path; for graphics, their display order in the page can be adjusted according to their hierarchical information. This dynamic adjustment method based on element display features improves the flexibility and adaptability of page conversion.
[0099] Step 104: Determine text input information of a large language model based on the page display information.
[0100] By converting page display information into input data for a large language model, the powerful semantic understanding capabilities of the LLM (Large Language Model) can be used to intelligently classify and semantically analyze page elements. This step ensures that the semantic information of page elements can be accurately identified and extracted, providing a semantic basis for subsequent grouping and layout optimization. By introducing LLM, the system can automatically identify the semantic categories of elements such as titles, text, and images, avoiding the limitations of traditional methods that rely on manual adjustments or fixed templates, and improving the automation and accuracy of conversion.
[0101] In addition, before step 104, the page display information may be converted into a JSON format file, and step 104 includes: extracting text information, image descriptions and context descriptions associated with page elements from the JSON format file obtained by converting the page display information as text input information for the large language model.
[0102] The JSON format can distribute the page display information in a structured manner, and extract text information, image descriptions, context descriptions and other content from the JSON format file. With the help of its data structured distribution characteristics, it can obtain the target content more intuitively and accurately, reducing the difficulty of obtaining text content.
[0103] In addition, for the image in the page element, the semantic description information of the image is generated based on the image description model. Among them, the image caption model is a machine learning model for generating semantic descriptions of images. Its main function is to generate a text description by analyzing the image content to summarize the main objects, scenes or activities in the image. In this technical solution, the image description model is used to generate semantic description information (image_desc) for the picture elements in the slide for subsequent semantic analysis and grouping processing.
[0104] On this basis, step 104 also includes: adding the semantic description information of the image to the text input information of the large language model. The semantic description information of the image (image_desc) provides a textual expression of the image content, which can help the large language model better understand the semantics of the image. For example, a picture showing a city landscape may generate a description "a picture showing a city landscape", which can help the model identify the relationship between the picture and the page theme. For another example, for an image of the butterfly life cycle, the image description model generates a semantic description of "a picture showing the butterfly life cycle", and the description is input into the large language model together with the text information, thereby ensuring the accurate grouping of the image on the semantic theme.
[0105] In summary, by combining the semantic description information of the image with the text content, annotation information, etc., the large language model can simultaneously process the multimodal information of text and images, thereby more comprehensively understanding the semantic structure of the page content. This multimodal information fusion improves the accuracy of semantic analysis and ensures that the model can more accurately classify and group page elements.
[0106] Step 106, input the text input information and a preset prompt for the text input information into the large language model, so that the large language model outputs page element classification information for grouping page elements according to semantic themes, wherein page elements with the same semantic theme are in the same group.
[0107] The page element classification information output by the large language model includes: the label of the page element and the functional type in the target slide. Through the semantic analysis of the large language model, the page elements can be automatically and intelligently classified to generate grouping labels for the page elements. This step solves the problem of lack of intelligent semantic analysis in the prior art, so that the semantic relationship of page elements can be accurately identified and utilized. For example, the relative positions of different contents such as titles, texts, and auxiliary elements can be intelligently adjusted through semantic analysis to ensure the logic and visual hierarchy of the content layout. Through this step, the system can automatically generate a page layout that conforms to semantic logic, reducing the workload of manual adjustment and improving conversion efficiency and layout quality.
[0108] Specifically, the text input information and preset prompts can be input into the large language model, which analyzes the semantic associations of page elements and outputs theme-based grouping information, such as grouping the title of "Insect Classification", related text, and insect pictures into a group of "Insect Theme". Through the semantic analysis capabilities of the large language model, the semantic relationships of page elements can be automatically identified and the elements can be grouped according to semantic themes. This intelligent grouping method avoids the limitations of traditional methods that rely on manual adjustments or fixed templates, and improves the automation and accuracy of the conversion.
[0109] At the same time, the semantic relationship between page elements and their visual priority are often not identified in related technologies. For example, the relative positions of different contents such as titles, texts, and auxiliary elements cannot be intelligently adjusted, resulting in a lack of logic and visual hierarchy in the content layout, affecting the overall display effect. In this regard, in this application, semantic grouping can be used to identify the logical relationship between page elements, ensuring that elements with the same semantic theme are reasonably organized together. For example, a title and its corresponding text content will be assigned to the same group to ensure that they maintain logical coherence in the converted page. This semantic-based grouping method makes the page layout more logical, enhances the visual hierarchy of the content, and improves the user's reading experience.
[0110] In addition, in order to achieve a reasonable layout in the related art, users are often required to manually adjust the position and size of each element, which is cumbersome and time-consuming, especially when the page elements are complex. In this regard, in this application, semantic grouping is directly performed through a large language model, which can automatically complete the classification and layout adjustment of elements, reduce manual intervention, and improve conversion efficiency. At the same time, compared with fixed templates, this dynamic grouping method can better handle complex content scenarios, ensure that the conversion results meet expectations, and can better adapt to actual scenarios with diverse slide content structures.
[0111] In summary, by introducing a large language model, intelligent semantic analysis and grouping of page elements can be achieved, which solves the problem of lack of intelligent semantic analysis in the existing technology. It not only improves the automation and accuracy of slide transitions, but also enhances the logic and visual hierarchy of page layout, reduces the workload of manual adjustment, and supports subsequent layout optimization, ultimately achieving efficient and intelligent horizontal to vertical conversion of slides.
[0112] In addition, after the large language model outputs the page element classification information, the page element classification information needs to be standardized, wherein the standardization includes deleting redundant tags, redundant functional types, and formatting symbols. Data standardization can delete redundant content in the page element classification information, reduce the difficulty of semantic analysis by the large language model, and help improve the accuracy of the output results of the large language model.
[0113] Step 108: Determine the size of the target shape page based on the page display information and a predetermined page size adjustment rule.
[0114] By dynamically calculating the size of the target page based on the page display information and the predetermined size adjustment rules, the page content can be adapted to the new display environment after conversion. This step solves the problem of reduced content readability caused by simple scaling methods in the prior art. By intelligently adjusting the page size, the text and images can remain clear and readable after conversion. In addition, dynamic adjustment of the page size can also avoid layout mismatch problems caused by fixed templates, improving the flexibility and adaptability of the conversion.
[0115] Specifically, first, determine a first ratio of the width of the current shape page to the width of the target shape page, and determine a second ratio of the height of the current shape page to the height of the target shape page, determine the minimum value of the first ratio and the second ratio as the lower limit value of the scaling ratio for converting the current shape page to the target shape page, and set the upper limit value of the scaling ratio to 1.
[0116] By calculating the ratio of the width and height of the current page to the target page, the system can determine the reasonable range of the zoom ratio. The minimum value of the first ratio (width ratio) and the second ratio (height ratio) is used as the lower limit of the zoom ratio to ensure that the page content will not be over-reduced and affect readability after zooming. At the same time, the upper limit is set to 1 to ensure that the page content will not exceed the display range of the target page due to over-enlargement. This step provides a reasonable range constraint for the subsequent zoom ratio search and avoids layout problems caused by improper zoom ratio.
[0117] Furthermore, within the zoom ratio range from the zoom lower limit value to the zoom upper limit value, a maximum feasible zoom ratio that satisfies a preset layout rule is searched by a binary search method, wherein the preset layout rule indicates that all page elements within the current shape page can be successfully displayed in the target shape page without changing their size.
[0118] By performing a binary search between the lower and upper scaling limits, the system can efficiently find the maximum feasible scaling ratio that meets the preset layout rules. The binary search algorithm quickly locates the optimal scaling ratio by continuously narrowing the search range, ensuring that the page content can be smoothly displayed on the target page without changing the size of the elements after scaling. This step solves the problem of unreasonable layout caused by fixed scaling ratios in traditional methods. By dynamically adjusting the scaling ratio, the system can generate the optimal layout solution based on the complexity of the page content and the size requirements of the target page.
[0119] Finally, the size of the target shape page is determined based on the size of the current shape page and the maximum feasible scaling ratio. This step ensures that the page content can adapt to the new display environment after conversion while maintaining the readability and visual integrity of the content. By dynamically adjusting the page size, the system can avoid layout confusion caused by fixed templates or simple scaling, and improve the display effect and user experience of the conversion results.
[0120] For example, the current shape page size is 1920x1080, and the target shape page size is 1080x1920. Calculate the width ratio 1080 / 1920=0.5625, the height ratio 1920 / 1080≈1.7778, take the minimum value 0.5625 as the lower limit of the scaling ratio, and the upper limit is 1. In the range of 0.5625 to 1, find the maximum feasible scaling ratio (such as 0.8) through binary search to meet the preset layout rules, that is, all page elements can be displayed in the target page when the size remains unchanged.
[0121] This technical solution determines the range of the zoom ratio by calculating the ratio of the width and height of the current page to the target page, and finds the maximum feasible zoom ratio that meets the preset layout rules through binary search. Through the above technical solution, the zoom ratio of the slide page can be intelligently determined to ensure that the page content can adapt to the size requirements of the target page after conversion. This solution solves the problem of unreasonable layout caused by fixed zoom ratio or simple zoom in traditional methods by calculating the reasonable range of zoom ratio and using binary search algorithm to efficiently find the maximum feasible zoom ratio. By dynamically adjusting the page size, the system can generate a logical and beautiful vertical page layout, improving the display effect and user experience.
[0122] Step 110 , based on the size of the target shape page, the page element classification information and a predetermined page element arrangement rule, determine arrangement information of the page elements in the current shape page within the target shape page.
[0123] By combining the target page size, element classification information and arrangement rules, the layout of page elements can be intelligently adjusted to ensure that the content maintains logic and visual hierarchy after conversion. This step solves the problem of insufficient flexibility in preset template conversion in the prior art. Through dynamically generated content containers and semantic analysis technology, the system can flexibly adjust the page layout according to the complexity of the content to adapt to diverse display needs. For example, elements such as titles, texts, and pictures can be reasonably arranged according to their semantic priority and size, avoiding the unreasonable layout problems caused by fixed templates in traditional methods. Through this step, the system can generate a logical and beautiful vertical page layout, which improves the display effect and user experience.
[0124] Specifically, first, the page layout parameters of the target shape page are initialized, wherein the page layout parameters include page margins, row spacing, column spacing, and element arrangement direction. For example, the page margins are set to 5%-10% of the page width, and the row spacing and column spacing are dynamically adjusted according to the number of elements, such as being set within the range of 10-20 pixels, and the element arrangement direction is from top to bottom and from left to right by default.
[0125] Next, all the page elements are traversed according to the element arrangement direction. During the traversal process, the priorities of all the page elements are obtained, and page space is preferentially allocated to page elements with higher priorities when arranging the elements.
[0126] Furthermore, if the cumulative width of the elements in the current row exceeds the width limit of the page margin, a line break operation is triggered; when performing the line break operation, the remaining arrangeable height of the target shape page is determined based on the page margin and the cumulative height of the elements. If the remaining arrangeable height is less than the height of the page elements to be arranged, a paging operation is triggered, or a trimming operation is performed on the page elements to be arranged.
[0127] For example, for a textured background element, its position in the current shape page is (0,0) to (1920,1080). It is scaled synchronously according to the maximum possible scaling ratio of 0.8, and the new size is 1536x864. If it is a non-textured background, the individual scaling ratio (such as 0.5625) is calculated based on the current page and the target page size, and the lower right corner coordinates are adjusted to (1080,1440). If the upper left corner exceeds the page, it is cropped to the center area.
[0128] In summary, the embodiments of the present application provide an intelligent method for converting slides from horizontal to vertical or from vertical to horizontal, which solves the problems of low manual adjustment efficiency, reduced readability due to simple scaling, insufficient flexibility of preset templates, and lack of intelligent semantic analysis in the prior art. By introducing a large language model and dynamic layout adjustment technology, the goals of automatically identifying the semantic relationship of page elements and intelligently adjusting the page layout can be achieved, ensuring that the content of the converted page is evenly distributed, logically clear, and visually beautiful, achieving automatic and efficient adaptive layout, high conversion efficiency, and high-quality conversion effects.
[0129] In addition, after step 110, the arrangement information may also be converted into a JSON format file.
[0130] It should be supplemented that, if the background element in the page element is a texture background or a gradient background, the background element is synchronously scaled according to the maximum feasible scaling ratio; if the background element in the page element is not a texture background or a gradient background, the individual scaling ratio of the background element is determined based on the respective sizes of the current shape page and the target shape page and the position of the background element, wherein the position of the background element includes the coordinates of the upper left corner vertex and the lower right corner vertex of the background element; the individual scaling ratio of the background element is consistent with the height scaling ratio or the width scaling ratio of the current shape page; based on the The lower right corner vertex coordinates of the background element in the current shape page are used to determine the lower right corner vertex coordinates of the background element in the target shape page; the upper left corner vertex coordinates of the background element in the target shape page are determined according to the individual scaling ratio of the background element and the lower right corner vertex coordinates of the background element in the target shape page; if the upper left corner vertex coordinates extend outside the target shape page, a cropping operation is performed on the background element so that the upper left corner vertex coordinates of the cropped background element are located within the target shape page, wherein the cropped background element is the image center area of the background element before cropping.
[0131] Figure 2 A flow chart of a method for switching slide presentation direction according to another embodiment of the present application is shown.
[0132] like Figure 2 As shown, according to another embodiment of the present application, a method for switching slide presentation direction includes:
[0133] Step 202: for the PPT file to be processed, parse its OfficeOpenXML structure and extract the page content, and call the model to generate the image description field for the image type element. Convert the resolution, element attributes (including position, size, type, image description, etc.) and hierarchical information of the page into a JSON format file. This step generates structured page data to provide input for subsequent processing.
[0134] This step parses the OfficeOpenXML structure of the PPT file to be processed and extracts the page content, including the resolution, element attributes (such as position, size, type, etc.) and hierarchical information of the page, and generates the image description (image_desc) field for the image element. The extracted and generated information is then converted into a JSON format file to synthesize structured page data to provide input for subsequent processing.
[0135] The process of parsing a PPT file may include the following specific operations: First, by reading the XML file <p:sldsz>The tag extracts the page size information, and uses the cx and cy attributes to represent the width and height of the slide page, thereby determining the resolution of the page; secondly, it extracts the page element information. For each page, locate <p:csld>The various page elements defined in the tag and its subtags. For a text box, <a:t>The text content is extracted from the tag, and its location, size and other attributes are obtained at the same time; for pictures, the location information and the associated image resource path are extracted; for graphics and other objects, their coordinates, hierarchy and size information in the page are extracted. In addition, the hierarchy information of page elements is determined by the definition order of the elements or the z-index related attribute values, which provides support for the rearrangement of content in subsequent layout optimization. While extracting page elements, the slide annotation content can also be parsed. The annotation content is usually stored in the XML file. <pp:notes>The tag contains <pp:txbody>Tags and their subtags <a:t>By extracting these annotation contents and combining them with the text and position information of page elements, additional context support can be provided for subsequent semantic analysis and grouping processing.
[0136] When extracting page element information, for elements of the image type, the image description (imagecaption) model is called to generate the semantic description information (image_desc) of the image, such as "an image describing terrestrial insects" or "an image showing a city landscape". The generated image description is stored in a JSON file as an attribute of the image element along with other information (such as position and size).
[0137] In the process of extracting information, each page element is defined as an independent data object, including its location, size, content, type, image description (if applicable) and other attributes, while recording its logical relationship in the page. For example, the text content in the text box and its font, font size, color and other style information will be fully extracted for subsequent semantic analysis and style adjustment.
[0138] In addition, to maintain the integrity of the page content as much as possible, while extracting the page resolution and element information, the background information of the page is also recorded, such as the background color, background image, or gradient fill effect. For pages without an explicitly defined background, the background is set to a blank background by default to provide integrity and consistency of the infrastructure data.
[0139] Step 204, read the page element information in the JSON file, combine the text and notes of the page, generate semantic grouping labels for each element through semantic analysis technology (using the LLM model), identify different semantic categories such as title, text, image, etc. in the page, and establish logical relationships between groups.
[0140] This step reads the page element information in the JSON file, combines the text content and remark information of the page, and uses the large language model (LLM) to generate semantic grouping tags for the page elements, identify different semantic types such as title, text, image, etc. in the page, and establish logical relationships between elements. In the specific implementation, the text content, remark information, position, size, type and image description information (image_desc) of the page elements are first extracted from the JSON file. These data provide complete input content for the semantic analysis model, strive to generate accurate semantic grouping results, and lay the foundation for subsequent logical relationship analysis and page content optimization. The processing of step 204 generally includes the following three logical stages:
[0141] The first logical phase is the semantic data construction phase. Specifically, the text information, remark content and geometric attributes of the elements in the page are extracted, combined with the context description, to construct the input data for semantic analysis.
[0142] The second logical stage is the semantic analysis model calling stage. Specifically, the large language model (LLM) is called, and the text information, image description and context description of the page elements are input according to the preset prompt template to complete the grouping and classification tasks of the page elements. The results returned by the model include the grouping label and preliminary semantic classification of each element.
[0143] An example of a large language model prompt template in a specific embodiment is as follows, wherein the {eles} and {notes} tags are replaced by the corresponding tag content of the JSON file generated in step 202 in actual implementation:
[0144] "The elements defined in PPT can be divided into the following according to their functions:
[0145] - Background image: The element type is a picture, and it usually occupies the entire page or most of the page, and is only used to provide a visual background;
[0146] -PPT limit box: The element type is a shape, and it usually occupies the entire page or most of the page. It is used to limit the layout size of all themes and text content in the PPT;
[0147] -Heading: Usually located at the top of the page, it summarizes the main content of the page;
[0148] - Combo box: no text content, the element type is shape, but the shape covers multiple other elements, and is used to confine multiple elements of the same theme to a geometric area;
[0149] - Subject content: contains specific information and data that explains the subject of the page in detail;
[0150] -Auxiliary elements: Elements used to enhance visual effects, such as icons, small pictures, etc., which are not related to the specific theme.
[0151] It is known that the original size of the PPT is: 960x540, and the elements contained in the page are as follows:
[0152] -----{eles}-----
[0153] Where id is the element id, name is the name of the element, type is the element type, position is the element position, size is the element size, and text is the text content of the element. For image type elements, image_desc is the image description recognized by the image-caption model.
[0154] The PPT notes are as follows:
[0155] -----{notes}-----
[0156] Please follow the steps below:
[0157] 1. Combine the PPT annotation content and the text of each element to analyze the multiple content themes expressed in the PPT;
[0158] 2. Analyze the function of each element in the page layout based on the page size, geometric position, size, and type of each element;
[0159] 3. For elements whose functions are thematic content, analyze the theme to which the element belongs and mark the grouping for each element;
[0160] 4. After the line break, output the final result in JSON format starting with `Results:`. The example is as follows:
[0161] Results:[{"id":"0","group":"title"},{"id":"1","group":"frame"},{"id":"2","group":"background image"},{"id":"3","group":"xx theme"},{"id":"4","group":"yy theme"},{"id":"5","group":"auxiliary elements"}].
[0162] Do not output anything else after this."
[0163] The third logical stage is the grouping result processing stage. Specifically, the grouping results returned by the semantic analysis model are cleaned and normalized, including deleting redundant tags through keyword matching, removing unnecessary group names and formatting symbols, and the generated grouping labels can accurately reflect the basic semantic classification of elements, providing support for subsequent logical analysis and optimization operations.
[0164] After the above logic stage processing, each page element is assigned a corresponding semantic grouping label. The result is presented in the JSON format defined by the prompt template (the Results section in the Prompt template above). For example, a title element may be grouped as "title", the text element below it is grouped as "subject content", and the image elements on the page may be grouped as "auxiliary elements" or "background images". Each element is identified by a unique id and assigned to a specific semantic grouping category.
[0165] The semantic description information of images plays a key role in grouping. For example, when they are associated with the theme of a paragraph of text content, they may be assigned to a grouping label consistent with the theme (such as "theme content"); when they are only used to enhance the visual effect or as a background, they may be assigned to "auxiliary elements" or "background images". The generated grouping results provide basic support for semantic analysis and layout adjustment of page content, and provide input data for logical analysis.
[0166] Step 206: Combined with the target resolution of the vertical page, a binary search algorithm is used to calculate the overall zoom ratio of the page and the offset of each element. To maintain the logical hierarchy of the page content, the relative positions of the elements in the group are adjusted to generate a preliminary content container.
[0167] This step calculates the overall scaling ratio of the page and the offset of each element based on the target size and ratio of the vertical page. Maintaining the logical hierarchy of the page content, the relative positions of elements in the same group are adjusted to generate a preliminary content container.
[0168] In the specific implementation, the calculation of the scaling ratio adopts an optimization method based on binary search to determine the maximum scaling ratio applicable to the target vertical page. In the initial stage, the upper and lower bounds of the scaling ratio are calculated, where the lower bound is determined by the minimum scaling factor calculated by the ratio of the page width and height. During the optimization process of the scaling ratio, the intermediate scaling ratio obtained each time is verified by continuously adjusting the binary search interval.
[0169] Specifically, for each intermediate zoom ratio value, a page re-layout operation is performed to detect whether the page content can be reasonably distributed within the target vertical page range under this ratio. If the current ratio is feasible, try to continue to increase the zoom factor; if not, reduce the factor range. During the search process, the previous minimum feasible zoom ratio is also retained to ensure that a better feasible solution can be returned during the zoom adjustment process. For example, assume that the size of the target page is 1080 pixels wide and 1440 pixels high, and the initial size of the root container is 1920 pixels wide and 1080 pixels high. In the binary search process, the initial range of the zoom ratio is first determined: the lower limit is the smaller value of the ratio of the target page width to the root container width and the ratio of the target page height to the root container height; the upper limit is 1, indicating no zoom. Next, by gradually adjusting the zoom ratio in the middle of the range, test whether it can meet the layout requirements. In each step of the test, take the middle value of the current upper and lower limits as the test ratio to verify whether the ratio is feasible. If it is feasible, it means that the page content can still fit into the target page. In this case, the lower limit of the zoom ratio is adjusted to this intermediate value to further try a larger ratio. If it is not feasible, the upper limit of the zoom ratio is adjusted to this intermediate value to narrow the range. In this way, the range of the zoom ratio is continuously narrowed, and finally the maximum feasible zoom ratio is determined while meeting the accuracy requirements.
[0170] Step 208, arrange the elements in the content container row by row, and dynamically wrap the elements according to the page width and height to ensure that the content is evenly distributed and logically clear. Provide paging output or cropping solutions for the part whose cumulative row height exceeds the page height.
[0171] This step arranges the elements in the content container row by row. Dynamic adjustments are made according to the width and height limits of the page to ensure that the content is evenly distributed and the logic is clear. In the specific implementation, the page layout parameters are first initialized based on the content container and its element information generated in step 206. These parameters include the row spacing, column spacing and initial arrangement direction of the content (e.g., arrangement from top to bottom or from left to right) of the page. At the same time, the basic constraints for layout adjustment are determined according to the width and height limits of the target page.
[0172] During the layout adjustment process, each element in the content container is traversed and its arrangement position is calculated in turn. For each element, its own width and height are compared with the cumulative width of adjacent elements. When the cumulative width reaches or exceeds the width limit of the page, a line break operation is triggered to arrange subsequent elements to a new row and update the vertical offset of the next row. At each line break, the cumulative height of the current row needs to be recorded to ensure that the next row can be arranged at the appropriate vertical position. For example, when the page width is 1080 pixels, if the cumulative width of the current row is 1040 pixels and the width of the element to be added is 100 pixels, the line break operation will move the element to the starting position of the next row.
[0173] For the elements in each group, adjust their layout in sequence according to their size, position and arrangement requirements, following the row and column rules. For larger elements (such as full-width images or main titles), it is necessary to calculate their row and column positions on the page separately, and try to ensure that they do not overlap with other elements. It should be understood that larger elements refer to content that occupies significant space on the page, such as a banner image that takes up the entire width of the page or a large title with a high visual priority. These elements need to be positioned first to avoid affecting the normal arrangement of other surrounding content. When dealing with larger elements, priority should be given to allocating sufficient space, and the arrangement of adjacent elements should be adjusted as needed to keep the page content as compact and logical as possible.
[0174] During the dynamic arrangement process, the cumulative height and width of each row or column need to be continuously updated, and the page content should be evenly distributed in terms of row spacing and column spacing. For example, for multi-column text content on a page, the column spacing can be appropriately reduced to improve the space utilization of the page, but at the same time, it is necessary to ensure that the line spacing between text paragraphs is sufficient to maintain good readability. In addition, for the part where the cumulative row height exceeds the page height, solutions such as paging processing and content cropping are provided.
[0175] Step 210: extract the background elements in the page, calculate their scaling and position offset, and adjust the layout consistency between the background and the content. For textured or gradient backgrounds, use appropriate tiling or cropping techniques to ensure the visual integrity of the background.
[0176] This step extracts the background elements in the page, calculates their scaling and position offset, and ensures that the background elements maintain layout consistency with the content elements. For textured or gradient backgrounds, appropriate tiling or cropping techniques are used to maintain the visual integrity of the background.
[0177] In the specific implementation, first extract the size and position information of the background element in the original page, such as the vertex coordinates of the upper left corner and the lower right corner. Based on the size change between the target page and the original page, calculate the global scaling ratio and position offset of the background element. It should be understood that the scaling ratio of the background element is determined by comparing the width and height of the original page with the width and height of the target page respectively, while trying to keep the aspect ratio of the background element unchanged to avoid unexpected visual deformation. For example, if the size of the original page is 1920 pixels × 1080 pixels, and the size of the target page is 1080 pixels × 1440 pixels, the scaling ratio of the background element should be based on the adaptation of the width or height, while taking into account the overall layout requirements of the vertical page.
[0178] By calculating the position offset of the background element in the target page, its new coordinate position in the vertical page can be determined. The vertex position of the lower right corner of the background element is adjusted first to ensure its alignment in the new page. Then, the final position of the upper left corner is deduced based on the scaled size of the background element to complete the repositioning of the background element. For example, when the background element is located in the lower right corner of the original page, the coordinates of its lower right corner are (1920, 1080). After calculation, if the corresponding scaled coordinates of the lower right corner of the target page are (1080, 1440), the position of the upper left corner of the background element will also be adjusted accordingly to try to ensure that the correct position and size of the background element fit the target page.
[0179] During the adjustment process, different processing strategies are adopted according to different types of background elements. For textured or gradient backgrounds, synchronized scaling and offset methods need to be applied to maintain the visual continuity and consistency of the background as much as possible. If the background pattern is large in size, it can be cropped to fit the target page and retain the most important visual area of the background; if the background pattern is small in size, the blank area of the page can be filled by texture tiling or gradient extension to avoid blank or uneven visual effects.
[0180] Step 212: Integrate the adjusted content container and background elements into a unified page structure and convert them into a JSON format file, recording the type, position, size and hierarchical relationship of each element. The resulting JSON file supports subsequent display and analysis.
[0181] This step integrates the adjusted content container and background elements into a unified page structure and converts them into a JSON format file. The JSON file records the type, position, size, and hierarchical relationship of each element to support subsequent display and analysis.
[0182] Figure 3 A schematic diagram of a slide show direction switching device according to an embodiment of the present application is shown.
[0183] like Figure 3 As shown, according to an embodiment of the present application, a slide display direction switching device 300 includes: an input parsing module 301, a semantic analysis module 302, a size adjustment module 303, a layout optimization module 304, a background adjustment module 305 and a verification output module 306.
[0184] The input parsing module 301 receives and parses the PPT file, extracts the page content and converts it into JSON format. The semantic analysis module 302 uses LLM to perform semantic analysis, generates semantic tags for page elements and groups them. The size adjustment module 303 calculates the scaling ratio and offset of the elements according to the target page size, and adjusts the page structure. The layout optimization module 304 optimizes the content layout based on the row and column arrangement rules to maintain visual balance. The background adjustment module 305 scales and offsets the background elements to ensure that the background and content maintain style consistency. The verification output module 306 integrates the adjusted page structure and outputs it as a JSON format file, and performs data integrity verification.
[0185] The device uses any one of the solutions described in the above embodiments, and therefore has all the above technical effects, which will not be repeated here.
[0186] Another embodiment of the present application also provides a slide display direction switching device, including: a page display information extraction unit, used to extract page display information in a target slide, wherein the page display information is used to reflect the display characteristics presented by the current shape page and the page elements respectively; a text input information determination unit, used to determine the text input information of the large language model based on the page display information; an LLM processing unit, used to input the text input information and a preset prompt for the text input information into the large language model, so that the large language model outputs page element classification information for grouping page elements according to semantic themes, wherein page elements with the same semantic theme are in the same group; a size adjustment unit, used to determine the size of the target shape page based on the page display information and a predetermined page size adjustment rule; an element rearrangement unit, used to determine the arrangement information of the page elements in the current shape page within the target shape page based on the size of the target shape page, the page element classification information and a predetermined page element arrangement rule.
[0187] In one embodiment of the present application, optionally, the page display information includes page display features of the current shape page and element display features of page elements within the current shape page, wherein the page display features include: one or more of resolution, size, page elements and page annotations; the element display features of the page elements include: one or more of position, hierarchy, resource path, size, content, type and element description text.
[0188] In one embodiment of the present application, optionally, the page display information extraction unit includes: a first extraction unit for locating the OfficeOpenXML file of the target slide <p:sldsz>Tags; Get the <p:sldsz>The cx attribute and the cy attribute in the tag, wherein the cx attribute and the cy attribute represent the width and height of the current shape page of the target slide respectively; based on the cx attribute and the cy attribute, the size and resolution of the current shape page are determined; the second extraction unit is used to locate the OfficeOpenXML file of the target slide <p:csld>Label and Description <p:csld>Subtags of the tag; based on the <p:csld>Label and Description <p:csld>The subtag of the tag determines the page elements in the current shape page and the element types of the page elements, wherein the types include: one or more of text boxes, images, graphics, and backgrounds; the third extraction unit is used to locate the OfficeOpenXML file of the target slide <pp:notes>label; <pp:txbody>Label and Description <pp:txbody>The text information stored in each of the sub-tags of the tag is determined as the page annotation of the current shape page; the fourth extraction unit is used to locate the OfficeOpenXML file of the target slide <a:t>Tag; based on the element type of each of the page elements, in the <a:t>The element display feature corresponding to each of the page elements is extracted from the text content stored in the tag.
[0189] In one embodiment of the present application, optionally, the device also includes: an image description unit, which is used to generate semantic description information of the image in the page element based on an image description model; then the text input information determination unit is also used to: add the semantic description information of the image to the text input information of the large language model.
[0190] In one embodiment of the present application, optionally, the device also includes: a first format conversion unit, used to convert the page display information into a JSON format file before the text input information determination unit determines the text input information of the large language model; a second format conversion unit, used to convert the arrangement information into a JSON format file after determining the arrangement information of the page elements in the current shape page in the target shape page; the text input information determination unit is used to: extract text information, image description and context description associated with the page elements in the JSON format file obtained by converting the page display information as the text input information of the large language model.
[0191] In one embodiment of the present application, optionally, the page element classification information output by the large language model includes: the label of the page element and the functional type in the target slide; the device also includes: a standardization processing unit, which is used to standardize the page element classification information after the LLM processing unit causes the large language model to output the page element classification information that groups the page elements according to semantic themes, wherein the standardization processing includes deleting redundant labels, redundant functional types and formatting symbols.
[0192] In one embodiment of the present application, optionally, the size adjustment unit includes: determining a first ratio of the width of the current shape page to the width of the target shape page, and determining a second ratio of the height of the current shape page to the height of the target shape page; determining the minimum value of the first ratio and the second ratio as the lower limit value of the scaling ratio for converting the current shape page to the target shape page, and setting the upper limit value of the scaling ratio to 1; within the scaling ratio range from the lower limit value to the upper limit value, searching for the maximum feasible scaling ratio that satisfies the preset layout rules by binary search, wherein the preset layout rules indicate that the page elements in the current shape page can be smoothly displayed in the target shape page without changing the size; determining the size of the target shape page based on the size of the current shape page and the maximum feasible scaling ratio.
[0193] In one embodiment of the present application, optionally, the element rearrangement unit is used to: initialize the page layout parameters of the target shape page, wherein the page layout parameters include page margins, row spacing, column spacing, and element arrangement direction; traverse all the page elements according to the element arrangement direction, and during the traversal process, obtain the priorities of all the page elements, and preferentially allocate page space to page elements with higher priorities when arranging elements; if the cumulative width of the elements in the current row exceeds the width limit of the page margins, trigger a line break operation; when performing a line break operation, determine the remaining arrangeable height of the target shape page based on the page margins and the cumulative height of the elements, and if the remaining arrangeable height is less than the height of the page elements to be arranged, trigger a paging operation, or perform a trimming operation on the page elements to be arranged.
[0194] In one embodiment of the present application, optionally, the device further includes:
[0195] The background adjustment unit is used to synchronously scale the background element according to the maximum feasible scaling ratio if the background element in the page element is a texture background or a gradient background; if the background element in the page element is not a texture background or a gradient background, determine the individual scaling ratio of the background element based on the respective sizes of the current shape page and the target shape page and the position of the background element, wherein the position of the background element includes the coordinates of the upper left corner vertex and the lower right corner vertex of the background element; the individual scaling ratio of the background element is consistent with the height scaling ratio or the width scaling ratio of the current shape page; based The lower right corner vertex coordinates of the background element in the current shape page are used to determine the lower right corner vertex coordinates of the background element in the target shape page; the upper left corner vertex coordinates of the background element in the target shape page are determined according to the individual scaling ratio of the background element and the lower right corner vertex coordinates of the background element in the target shape page; if the upper left corner vertex coordinates extend outside the target shape page, a cropping operation is performed on the background element so that the upper left corner vertex coordinates of the cropped background element are located within the target shape page, wherein the cropped background element is the image center area of the background element before cropping.
[0196] The device uses any one of the solutions described in the above embodiments, and therefore has all the above technical effects, which will not be repeated here.
[0197] Figure 4 A block diagram of a computer device according to an embodiment of the present application is shown.
[0198] like Figure 4 As shown, a computer device according to an embodiment of the present application includes: a processor, an internal memory, a non-volatile memory (external memory), an operating system, a system bus, and an I / O interface. The non-volatile memory stores an operating system and a slide horizontal to vertical conversion program. The processor drives the slide horizontal to vertical conversion program by calling program instructions in the internal memory to implement the technical solution described in any of the above embodiments, and each module exchanges data through the system bus. The I / O interface connects the input device and the output device, and is used to receive slide files and output the adjusted vertical page content.
[0199] In addition, in one embodiment, the present application provides a computer device, which may be a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, the method described in any of the above embodiments can be implemented.
[0200] In one embodiment, the present application further provides a computer device, which may be a client, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, the method described in any of the above embodiments can be implemented.
[0201] Any of the above-mentioned computer devices in the embodiments of the present application may exist in various forms, including but not limited to:
[0202] (1) Mobile communication devices: These devices are characterized by their mobile communication functions and their main purpose is to provide voice and data communications. These terminals include: smart phones (such as iPhone), multimedia phones, functional phones, and low-end phones.
[0203] (2) Ultra-mobile personal computer devices: These devices fall into the category of personal computers, have computing and processing capabilities, and generally also have mobile Internet access features. These terminals include: PDA, MID and UMPC devices, such as iPad.
[0204] (3) Portable entertainment devices: These devices can display and play multimedia content. They include audio and video players (such as iPods), handheld game consoles, e-books, as well as smart toys, wearable devices, and portable car navigation devices.
[0205] (4) Server: A device that provides computing services. The server consists of a processor, hard disk, memory, system bus, etc. The server is similar to a general computer architecture, but because it needs to provide highly reliable services, it has higher requirements in terms of processing power, stability, reliability, security, scalability, and manageability.
[0206] (5) Other electronic devices with data interaction functions.
[0207] In addition, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to perform the following steps:
[0208] Extracting page display information from the target slide, wherein the page display information is used to reflect the display characteristics of the current shape page and the page elements;
[0209] Determining text input information of a large language model based on the page display information;
[0210] Inputting the text input information and a preset prompt for the text input information into the large language model, so that the large language model outputs page element classification information for grouping page elements according to semantic themes, wherein page elements with the same semantic theme are in the same group;
[0211] Determining the size of the target shape page based on the page display information and a predetermined page size adjustment rule;
[0212] Arrangement information of the page elements in the current shape page within the target shape page is determined based on the size of the target shape page, the page element classification information and a predetermined page element arrangement rule.
[0213] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant description in the aforementioned method embodiment. In order to avoid repetition, they will not be described one by one here.
[0214] The above describes the technical solution of the present application in detail in combination with the accompanying drawings. Through the technical solution of the present application, the problems of low manual adjustment efficiency, reduced readability due to simple scaling, insufficient flexibility of preset templates, and lack of intelligent semantic analysis in the prior art are solved. By introducing a large language model and dynamic layout adjustment technology, the goal of automatically identifying the semantic relationship of page elements and intelligently adjusting the page layout can be achieved, ensuring that the converted page content is evenly distributed, logically clear, and visually beautiful, achieving automatic and efficient adaptive layout, high conversion efficiency, and high-quality conversion effects.
[0215] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when determining" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0216] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms of "a", "said", and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0217] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0218] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0219] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0220] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.< / a:t> < / a:t> < / pp:txbody> < / pp:txbody> < / pp:notes> < / p:csld> < / p:csld> < / p:csld> < / p:csld> < / p:sldsz> < / p:sldsz> < / a:t> < / pp:txbody> < / pp:notes> < / a:t> < / p:csld> < / p:sldsz> < / a:t> < / a:t> < / a:t> < / pp:notes> < / pp:txbody> < / pp:txbody> < / pp:notes> < / p:csld> < / p:csld> < / p:csld> < / p:csld> < / p:csld> < / p:sldsz> < / p:sldsz> < / a:t> < / a:t> < / pp:txbody> < / pp:txbody> < / pp:notes> < / p:csld> < / p:csld> < / p:csld> < / p:csld> < / p:sldsz> < / p:sldsz> < / a:t> < / a:t> < / pp:txbody> < / pp:txbody> < / pp:notes> < / p:csld> < / p:csld> < / p:csld> < / p:csld> < / p:sldsz> < / p:sldsz>
Claims
1. A method for switching slide show direction, characterized in that: include: Extracting page display information from the target slide, wherein the page display information is used to reflect the display characteristics of the current shape page and the page elements; Determining text input information of a large language model based on the page display information; Inputting the text input information and a preset prompt for the text input information into the large language model, so that the large language model outputs page element classification information for grouping page elements according to semantic themes, wherein page elements with the same semantic theme are in the same group; Determining the size of the target shape page based on the page display information and a predetermined page size adjustment rule; Determining arrangement information of the page elements in the current shape page within the target shape page based on the size of the target shape page, the page element classification information and a predetermined page element arrangement rule; The determining the size of the target shape page based on the page display information and a predetermined page size adjustment rule includes: Determining a first ratio of a width of the current shape page to a width of the target shape page, and determining a second ratio of a height of the current shape page to a height of the target shape page; Determine the minimum value of the first ratio and the second ratio as the lower limit value of the scaling ratio for converting the current shape page into the target shape page, and set the upper limit value of the scaling ratio to 1; In the zoom ratio range from the zoom lower limit value to the zoom upper limit value, searching for the maximum feasible zoom ratio that satisfies the preset layout rule by binary search, wherein the preset layout rule indicates that all page elements in the current shape page can be successfully displayed in the target shape page without changing their size; Based on the size of the current shape page and the maximum feasible zoom ratio, the size of the target shape page is determined.
2. The method according to claim 1, characterized in that The page display information includes the page display characteristics of the current shape page and the element display characteristics of the page elements in the current shape page, wherein The page display characteristics include: one or more of resolution, size, page elements and page annotations; The element display characteristics of the page element include: one or more of: position, level, resource path, size, content, type and element description text.
3. The method according to claim 2, characterized in that The step of extracting the display information of the current shape page in the target slide includes: Locate the OfficeOpenXML file of the target slide <p:sldsz> Label;< / p:sldsz> Get the <p:sldsz> The cx attribute and the cy attribute in the tag, wherein the cx attribute and the cy attribute represent the width and height of the current shape page of the target slide respectively;< / p:sldsz> Determining the size and resolution of the current shape page based on the cx attribute and the cy attribute; and Locate the OfficeOpenXML file of the target slide <p:csld>Label and Description <p:csld> Subtags of a tag;< / p:csld> < / p:csld> Based on the <p:csld>Label and Description <p:csld> The sub-tag of the tag determines the page elements in the current shape page and the element types of the page elements, wherein the types include: one or more of text box, image, graphic, and background;< / p:csld> < / p:csld> Locate the OfficeOpenXML file of the target slide <pp:notes> Label;< / pp:notes> The <pp:txbody>Label and Description <pp:txbody> The text information stored in each sub-tag of the tag is determined as the page annotation of the current shape page;< / pp:txbody> < / pp:txbody> Locate the OfficeOpenXML file of the target slide <a:t> Label;< / a:t> Based on the element type of each of the page elements, <a:t> The element display feature corresponding to each of the page elements is extracted from the text content stored in the tag.< / a:t> 4. The method according to claim 3, characterized in that Also includes: For the image in the page element, generating semantic description information of the image based on the image description model; Then, determining the text input information of the large language model based on the page display information includes: The semantic description information of the image is added to the text input information of the large language model.
5. The method according to claim 1, characterized in that: Before determining the text input information of the large language model based on the page display information, the method further includes: Convert the page display information into a JSON format file; Then, determining the text input information of the large language model based on the page display information includes: Extracting text information, image description and context description associated with page elements from the JSON format file obtained by converting the page display information as text input information of the large language model; After determining the arrangement information of the page elements in the current shape page in the target shape page, the method further includes: The arrangement information is converted into a JSON format file.
6. The method according to claim 5, characterized in that The page element classification information output by the large language model includes: the label of the page element and the function type in the target slide; After the large language model is made to output page element classification information for grouping page elements according to semantic topics, the method further includes: The page element classification information is standardized, wherein the standardization includes deleting redundant tags, redundant function types and formatting symbols.
7. The method according to claim 1, characterized in that The determining, based on the size of the target shape page, the page element classification information and a predetermined page element arrangement rule, arrangement information of the page elements in the current shape page within the target shape page comprises: Initializing page layout parameters of the target shape page, wherein the page layout parameters include page margins, row spacing, column spacing, and element arrangement direction; Traverse all the page elements according to the element arrangement direction, and during the traversal process, Obtaining the priorities of all the page elements, and allocating page space to page elements with higher priorities first when arranging the elements; If the accumulated width of the elements in the current row exceeds the width limit of the page margin, a line break operation is triggered; When performing a line break operation, the remaining arrangeable height of the target shape page is determined based on the page margin and the cumulative height of the elements. If the remaining arrangeable height is less than the height of the page elements to be arranged, a paging operation is triggered, or a trimming operation is performed on the page elements to be arranged.
8. The method according to claim 7, characterized in that Also includes: If the background element in the page element is a texture background or a gradient background, the background element is synchronously scaled according to the maximum feasible scaling ratio; If the background element in the page element is not a texture background or a gradient background, then, based on the respective sizes of the current shape page and the target shape page and the position of the background element, the individual scaling ratio of the background element is determined, wherein: The position of the background element includes the coordinates of the upper left corner vertex and the lower right corner vertex of the background element; The individual scaling ratio of the background element is consistent with the height scaling ratio or the width scaling ratio of the current shape page; Determine the coordinates of the lower right corner vertex of the background element in the target shape page based on the coordinates of the lower right corner vertex of the background element in the current shape page; Determine the upper left corner vertex coordinates of the background element in the target shape page according to the individual scaling ratio of the background element and the lower right corner vertex coordinates of the background element in the target shape page; If the upper left corner vertex coordinates extend outside the target shape page, a cropping operation is performed on the background element so that the upper left corner vertex coordinates of the cropped background element are located within the target shape page, wherein the cropped background element is the image center area of the background element before cropping.
9. A slide show direction switching device, characterized in that: include: A page display information extraction unit, used to extract page display information from a target slide, wherein the page display information is used to reflect the display characteristics of the current shape page and the page elements; A text input information determining unit, configured to determine text input information of a large language model based on the page display information; An LLM processing unit is used to input the text input information and a preset prompt for the text input information into the large language model, so that the large language model outputs page element classification information for grouping page elements according to semantic themes, wherein page elements with the same semantic theme are in the same group; A size adjustment unit, configured to determine the size of the target shape page based on the page display information and a predetermined page size adjustment rule; An element rearrangement unit, configured to determine arrangement information of the page elements in the current shape page within the target shape page based on the size of the target shape page, the page element classification information and a predetermined page element arrangement rule; The size adjustment unit is used to: determine a first ratio of the width of the current shape page to the width of the target shape page, and determine a second ratio of the height of the current shape page to the height of the target shape page; determine the minimum value of the first ratio and the second ratio as the lower limit value of the scaling ratio for converting the current shape page to the target shape page, and set the upper limit value of the scaling ratio to 1; within the scaling ratio range from the lower limit value to the upper limit value, search for the maximum feasible scaling ratio that meets the preset layout rules by binary search, wherein the preset layout rules indicate that the page elements in the current shape page can be smoothly displayed in the target shape page without changing the size; determine the size of the target shape page based on the size of the current shape page and the maximum feasible scaling ratio.
Citation Information
Patent Citations
Method for use in processing text and voice information, and terminal
CN108885614A
Online webpage generation method and device based on document processing
CN119003906A