Automatic figure arrangement method, device and equipment of presentation and storage medium
By extracting multiple description sentences from presentations and using a cross-modal retrieval model to automatically screen images, the problem of low image matching efficiency in existing technologies is solved, and an efficient and accurate automatic image matching process is achieved.
Patent Information
- Application Number
- CN202410564728.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-08
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-05-08
AI Technical Summary
Existing presentation editing software is inefficient in terms of image matching, especially when it is difficult to quickly find suitable illustrations among massive image resources, resulting in high labor costs.
Automatic image matching is achieved by extracting multiple description sentences from the page to be matched with images in the presentation, using a cross-modal retrieval model to match candidate images, and screening out target images based on the matching degree.
It improves the efficiency and accuracy of presentation illustrations, reduces manual intervention, and improves the recall rate and quality of illustrations.
Smart Images

Figure CN118537446B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to technical fields such as artificial intelligence, deep learning, and natural language understanding. Background Art
[0002] Presentations clearly and concisely convey the main idea and core content. Presentations are commonly used in teaching, research, and business activities. Compared to text-only presentations, presentations with both text and images can convey the core idea more vividly. Therefore, creating a presentation requires not only organizing the text but also providing relevant images. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, device and storage medium for automatically matching images for a presentation.
[0004] According to one aspect of the present disclosure, a method for automatically adding images to a presentation is provided, comprising:
[0005] Extract multiple description sentences from the page to be matched with pictures in the presentation;
[0006] Match each description statement with the corresponding candidate image to obtain a candidate image set;
[0007] Determine the matching degree between each candidate image in the candidate image set and the page to be matched;
[0008] Based on the matching degree between each candidate image and the page to be matched, a target image is selected for the page to be matched as an illustration of the page to be matched.
[0009] According to another aspect of the present disclosure, there is provided an automatic image matching device for a presentation, comprising:
[0010] An extraction module is used to extract multiple description sentences from the page to be matched with pictures in the presentation;
[0011] A matching module is used to match the corresponding candidate images for each description statement to obtain a candidate image set;
[0012] A determination module is used to determine the matching degree between each candidate image in the candidate image set and the page to be matched;
[0013] The screening module is used to screen out target images for the page to be matched based on the matching degree between each candidate image and the page to be matched, so as to serve as illustrations for the page to be matched.
[0014] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.
[0018] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.
[0019] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.
[0020] In the disclosed embodiment, by extracting multiple description sentences, suitable illustrations can be automatically and accurately selected for the page to be illustrated.
[0021] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0023] Figure 1 is a flowchart of a method for automatically assigning images to a presentation according to an embodiment of the present disclosure;
[0024] Figure 2 is a schematic diagram of generating multiple description sentences based on the title and the body according to an embodiment of the present disclosure;
[0025] Figure 3 is a schematic diagram of a process for further limiting the requirements for candidate images according to an embodiment of the present disclosure;
[0026] Figure 4 is a schematic diagram of extracting various description statements according to an embodiment of the present disclosure;
[0027] Figure 5 is a schematic diagram of searching for candidate images and final illustrations using a cross-modal retrieval model according to an embodiment of the present disclosure;
[0028] Figure 6 is a schematic diagram of a process for obtaining candidate images for various description sentences according to an embodiment of the present disclosure;
[0029] Figure 7is a flowchart of determining matching degrees of each candidate image in a candidate image set to a page to be matched according to an embodiment of the present disclosure;
[0030] Figure 8 is a diagram of a comparison result according to an embodiment of the present disclosure;
[0031] Figure 9 is a structural diagram of an automatic image matching device for a presentation according to an embodiment of the present disclosure;
[0032] Figure 10 is a block diagram of an electronic device for implementing an automatic image matching method for a presentation according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0033] Exemplary embodiments of the present disclosure are described herein with reference to the accompanying drawings, which are included to provide a thorough understanding of embodiments of the present disclosure by a person of ordinary skill in the art, and should not be construed as limiting the present disclosure to particular embodiments. Accordingly, those of ordinary skill in the art will recognize that there are various modifications and changes that can be made thereto without departing from the scope of the present disclosure. Also, descriptions of well-known functions and constructions are omitted for clarity and conciseness.
[0034] The terms "first", "second", and the like in the present disclosure are used to distinguish similar objects, and do not necessarily have to be used to describe a particular order or sequence. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, inclusion of a series of steps or units. The method, system, product, or device does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the processes, methods, products, or devices.
[0035] The presentation in the embodiments of the present disclosure can include PPT (Power Point, software for making slides and reports), word (Office Word, document), WPS (Word Processing System, word processing system), PDF (Portable Document Format, portable document format), and the like, editable electronic documents of graphics and texts.
[0036] Although some editing software of presentations can provide very rich functions to help users edit presentations with graphics and texts, the manual editing method is low in efficiency. Especially in terms of image matching, how to find suitable images in the face of massive image resources also needs to consume time and labor costs.
[0037] In view of this, in order to improve the editing efficiency of presentations, the embodiment of the present disclosure provides a method for automatically adding pictures to presentations. It can be understood that the embodiment of the present disclosure processes any page of a presentation to be added with pictures in the same way, and the embodiment is described with any page to be added with pictures. After the text information is edited on the page to be added with pictures, automatic picture addition can be achieved based on the method in the embodiment of the present disclosure. Figure 1 The flowchart of the method is shown in FIG. , which includes the following contents:
[0038] S101, extracting multiple description sentences from a page of a presentation to be matched with pictures.
[0039] Each description statement extracted in the disclosed embodiment can describe the core content of the page to be illustrated from different dimensions or granularity, expressing the page's requirements for illustrations. Therefore, the generated multiple description statements express similar illustration requirements, but in different ways. The extraction of multiple description statements will be described later, so we will not elaborate on them here.
[0040] S102: Match corresponding candidate images for each description statement to obtain a candidate image set.
[0041] For example, in the case where multiple description sentences include description sentence 1 and description sentence 2, description sentence 1 matches m candidate images, and description sentence 2 matches n candidate images, then the candidate image set is the union of the candidate images matched by description sentence 1 and description sentence 2. For example, the candidate image set may include (m+n) images.
[0042] S103: Determine the matching degree between each candidate image in the candidate image set and the page to be matched with an image.
[0043] S104 : Based on the matching degree between each candidate image and the page to be matched with an image, a target image is selected for the page to be matched with an image, and the target image is used as an illustration for the page to be matched with an image.
[0044] During implementation, a preset number of candidate images with high matching degrees are preferentially selected to illustrate the page to be illustrated. The preset number is determined based on the number of illustrations required for the page to be illustrated. In the embodiment of the present disclosure, the character content and typesetting style in the page to be illustrated can be generated based on manual operation, or can be generated based on the function of automatically generating text content. Whether it is manual or automatic, the page to be illustrated can correspond to the required number of illustrations. Therefore, when the illustrations are finally screened, the preset number of candidate images can be selected as illustrations in the order of the matching degree between each candidate image and the page to be illustrated from high to low.
[0045] In summary, in the embodiments of the present disclosure, by extracting multiple description statements, the requirements of the page to be illustrated for illustrations can be described from different dimensions and / or granularity. The multiple description statements can be used to recall candidate images related to the page to be illustrated as much as possible, thereby improving the recall rate of the candidate image set. In this way, suitable images are preliminarily screened out. On this basis, by determining the degree of match between each candidate image and the page to be illustrated, illustrations that can assist the page to be illustrated in expressing the core idea can be further screened out. The entire illustration process can automatically and step by step screen out suitable illustrations for the page to be illustrated without the need for manual intervention.
[0046] For ease of understanding, the process of generating multi-way description statements and filtering images is explained below.
[0047] 1) Generate multi-way description statements
[0048] In the disclosed embodiments, the requirements for the image page to be illustrated can be described from various dimensions and / or granularities. In one possible implementation, the title information and text content of the page to be illustrated are extracted from the presentation; and multiple description statements are generated based on the title information and text content.
[0049] Taking a PPT document as an example, a PPT document may include a title and a body text, and multiple descriptive sentences may be extracted based on these text contents.
[0050] For example, the title serves as an outline and summarizes the core content. The body of the presentation describes the core content in detail.
[0051] In the disclosed embodiment, the requirements for illustrations can be described at different granularities through the title and body content. The title and body content can concisely and clearly express the illustration requirements of the page to be illustrated. Moreover, the title and body content are easy to extract, and multiple description statements can be obtained accurately and quickly, thereby improving the efficiency of illustration for the page to be illustrated.
[0052] Furthermore, in order to improve the recall rate of the candidate image set, in the embodiment of the present disclosure, at least the following two operations may be performed to generate multiple description sentences based on the title and the text content, such as Figure 2 As shown:
[0053] S201, according to the level of the title, sequentially splice the titles of each level in the title information to obtain a first description sentence.
[0054] For example, Figure 2 As shown in the figure, the page to be matched with pictures is used to introduce the freehand description of mountains in ancient poems. Assuming that the title of the current page to be matched with pictures is from the first-level title Poems, to the second-level title Mountain, and then to the third-level title Mountain Flowers and Peach Blossoms. Then this page belongs to the first level of the multi-level title. Figure 2As shown, the title of the page to be matched with the picture only includes the third-level title "Mountain Flowers of Peach Blossoms", and the content of the page body includes:
[0055] The beauty of April fades, and the peach blossoms at the mountain temple begin to bloom. — Bai Juyi, "Peach Blossoms at Dalin Temple"
[0056] April is a clear and gentle month, and after a rain has just cleared up, the South Mountain in front of the window becomes clear. No more willow catkins are blown by the wind, only sunflowers lean towards the sun. — Sima Guang
[0057] To extract the title information of the page to be matched with pictures from the presentation, if the page to be matched with pictures includes multi-level titles, the title information (including multi-level titles) is extracted from the page to be matched with pictures; if the page to be matched with pictures only includes the title of the page itself, the title information (including multi-level titles) is extracted for the page to be matched with pictures from the title outline of the presentation.
[0058] The title outline can use different symbols to indicate the heading level. The following example uses different numbers of # symbols to indicate the heading level:
[0059] #Poetry(first level title)
[0060] ##Mountain (Secondary Title)
[0061] ###Mountain Flower Peach Sunflower (Level 3 Title)
[0062] like Figure 2 As shown, after extracting the title information, it is sorted and spliced according to the title level, and the first description sentence is obtained: Poetry Mountain Flower Peach Kui.
[0063] It should be noted that the order of splicing can be from high to low in title level, or from low to high in title level.
[0064] In addition, when splicing, titles of different levels can be distinguished by spaces or preset connectors, which is not limited in the embodiment of the present disclosure.
[0065] The multi-level titles spliced together through the first-line description statements can clearly express the relationship between the page to be matched with pictures and the previous text, thereby making it easier to sort out the relationship between the current page to be matched with pictures and the expressed title, and recall candidate images that can express the core framework.
[0066] S202: Concatenate the first description statement and the body content in a preset order to obtain a second description statement.
[0067] Based on Figure 2Taking the image page shown above as an example, the first-path description and the main text can be combined. The second-path description, through the combined title, can, to a certain extent, reflect the relationship between the content of the current image page and the core content of the previous text. It can also accurately describe the image content requirements of the current image page through the main text. Therefore, the second-path description can help recall candidate images that both express the overall concept and the detailed requirements.
[0068] S203: Based on the first keyword extraction method, extract at least one first-category keyword from the title information and the text content to construct a third description sentence.
[0069] Continue with Figure 2 For example, the page for images to be added can be used to extract the keyword "April" based on word frequency. The term "April" further limits the time period, placing requirements on image content from a temporal perspective.
[0070] There are many ways to extract keywords, such as extracting keywords based on word frequency, extracting keywords based on a neural network model, extracting keywords based on a candidate keyword list, extracting keywords based on semantics or part of speech, etc. The specific method of extracting keywords is not limited in the embodiments of the present disclosure. It should be noted that in the embodiments of the present disclosure, the first keyword extraction method is used to extract the first category of keywords to construct the third description statement. The requirements for the content of the accompanying pictures can be described from the perspective of keywords in order to improve the recall rate of candidate images.
[0071] S204, extracting target-type words in order from low to high levels of the titles; wherein, if target-type words are extracted from any level of title, target-type words will no longer be extracted from titles with a level higher than any level of title.
[0072] For example, assuming there are n levels of titles in total, if the target type words are extracted from the n-level title, the operation is terminated; if the target type words are not extracted, the target type words will continue to be extracted from the n-1-level title. If the target type words are extracted, the operation is terminated, otherwise the extraction will continue from the previous level (n-2) title.
[0073] S205: Construct a fourth description sentence based on the extracted target-type words.
[0074] For example, target words can be defined based on part of speech, such as noun, verb, etc. Target words can also be further defined based on what kind of noun they are, such as place names, personal names, animals, plants, etc.
[0075] In the disclosed embodiments, different extraction templates can be generated for different target terms. When automatically matching images, the user simply selects the corresponding template to complete the image search. For example, a template for introducing a person can be used to create images specifically for presentations about that person. A template for introducing a travel guide can be used to extract images of place names and scenery for matching images.
[0076] Continue with Figure 2 Taking the page of pictures to be matched in as an example, assuming that the target words are plant nouns, peach and sunflower can be further extracted to further limit the image content of the candidate images.
[0077] By extracting target-class terms, we can further constrain image details, improve image recall, and enhance the efficiency of automatic image matching. Furthermore, target-class terms are specifically set for titles. Starting extraction from the last-level title improves the efficiency of extracting target-class terms and, consequently, the efficiency of automatic image matching.
[0078] Of course, it should be noted that descriptions from different paths can overlap in content. For example, extracted keywords and target-category words can overlap. Different descriptions place different demands on images from different perspectives, which can improve image matching efficiency.
[0079] For the fourth-path description, if multiple target-category words are obtained, these words can be concatenated to form the fourth-path description. The concatenation can be performed sequentially according to the weight of each target-category word. Weights can be assigned based on word frequency, with higher-frequency words receiving higher weights.
[0080] In other embodiments, in order to improve the recall rate while limiting the number of recalled images as much as possible, constructing the fourth description statement can be implemented as follows: adding a pre-set formatted explanatory text to the target class words to obtain the fourth description statement.
[0081] Continue with Figure 2 For example, if the target word extracted is "peach", then it can be further interpreted as "peach blossoms in the mountains" based on the description of the page to be matched. This avoids recalling peach blossoms and peaches in the plains.
[0082] In addition, assuming that the target term is a person's name, when a person's name is extracted, it can be further explained that: *** is a person. This can help understand the type of object pointed to by the target term and avoid recalling images of flowers when the person's name contains the word "flower".
[0083] In summary, explanatory text can be defined based on templates, or it can be automatically generated using natural language understanding (NLU) based on the textual information on the image page. In practice, a neural network model can be trained to analyze the information on the image page and automatically generate appropriate explanatory text for the target category.
[0084] It can be seen that by interpreting the text, we can provide restrictive descriptions for target words, reduce misunderstandings of target words, improve the recall rate, and improve the recall accuracy of candidate images.
[0085] It should be noted that when multiple target-class terms are extracted, each target-class term and its corresponding explanatory text form a holistic expression, facilitating the identification of the correspondence between the target-class term and the explanatory text. During implementation, the explanatory text can be represented by a designated symbol. For example, the expression A[A is a person] can serve as the holistic expression for the target-class term A. Ultimately, the holistic expressions of the different target-class terms can be sequentially concatenated to form a fourth descriptive statement.
[0086] In the case of obtaining the above multiple description sentences, the recall rate can be effectively improved. However, in order to improve the quality of the candidate images, in the embodiment of the present disclosure, at least one description sentence can be optimized to further define the requirements for the candidate images. It can be implemented as follows Figure 3 As shown:
[0087] S301, extracting a target proper noun from a preset proper noun type set from the title information and the text content.
[0088] S302: When the target proper noun is extracted, the target proper noun is added to at least one description statement in the multiple description statements.
[0089] In the embodiment of the present disclosure, the type of the target proper noun extracted from the preset proper noun type set may be the same as or different from the type of the target class word, and can be determined according to actual needs during implementation.
[0090] If the two are the same, the target proper name can be added to the description statements except the fourth description statement. If the two are different, the target proper name can be added to each of the four description statements.
[0091] After adding the target proper noun, each description statement further defines the requirements for illustrations in a fine-grained manner, which can improve the quality of the recalled candidate images while increasing the recall rate.
[0092] For example, in a PPT presentation about a person, the title is represented by # on the current page to be matched with images. One # represents a first-level title, two ## represent a second-level title, and so on up to a fourth-level title. The content after the fourth-level title is the main text. The content of the page to be matched with images is as follows:
[0093] #All Table Tennis Grand Slam Players (First Level Title)
[0094] ##Grand Slam Players (Secondary Title)
[0095] ###AAA (third level title)
[0096] #### Men's Singles Grand Slam (Level 4 Title)
[0097] He won the men's singles title at the 2012 World Table Tennis Championships, the men's singles and team titles at the 2012 Olympic Games, and the men's singles title at the 2012 World Table Tennis Championships. (Text)
[0098] like Figure 4 As shown in the figure, the description sentences extracted based on the above method are:
[0099] The first description sentence constructed based on the title: All-time Table Tennis Grand Slam Players Grand Slam Players AAA Men's Singles Grand Slam;
[0100] The second description sentence constructed based on the title and body content: Past table tennis Grand Slam players, Grand Slam players, AAA men's singles Grand Slam, ** World Table Tennis Championships men's singles champion, ** Olympic table tennis men's singles and men's team champions, ** World Table Tennis Championships men's singles champion.
[0101] The third description sentence constructed based on the keywords in the title and body content: AAA Grand Slam Table Tennis Champion Men's Singles.
[0102] Based on the target words, the fourth description sentence constructed for the name in this example is: AAA [is a person].
[0103] Assuming that the extracted proper name is also a person's name, the target proper name to be extracted is AAA. Adding the target proper name AAA to the first three-way description statement, the final four-way description statement is:
[0104] The final first-way description sentence: All-time table tennis Grand Slam players Grand Slam players AAA men's singles Grand Slam AAA;
[0105] The final second description sentence: All-time table tennis Grand Slam players, Grand Slam players AAA men's singles Grand Slam, ** World Table Tennis Championships men's singles champion, ** Olympic table tennis men's singles and men's team champions, ** World Table Tennis Championships men's singles champion AAA.
[0106] The final third description sentence: AAA Grand Slam Table Tennis Champion Men's Singles.
[0107] Based on the target words, the fourth description sentence constructed for the name in this example is: AAA [is a person].
[0108] It is understandable that even if the extracted target proper noun is repeated in some road description sentences, it can still play a role in emphasizing the target proper noun and help screen out candidate images that are strongly related to the target proper noun.
[0109] In summary, we've introduced the process of constructing a multi-path description statement. The following section further explains how to match appropriate illustrations based on each description statement.
[0110] 2) Filter images
[0111] In the disclosed embodiment, candidate images and final illustrations may be found based on a cross-modal retrieval model.
[0112] like Figure 5 As shown, the key parts of the cross-modal retrieval model may include an image encoder and a text encoder.
[0113] Among them, the image encoder can adopt the VIT (Vision Transformer, visual processing model) model, and the text encoder can adopt the BERT (Bidirectional Encoder Representation from Transformers, pre-trained language representation model) model.
[0114] In the embodiment of the present disclosure, a training sample is constructed by an image-text pair, that is, an image and its corresponding text description serve as a training sample.
[0115] The retrieval model can be fine-tuned based on the CLIP model (Contrastive Language-Image Pre-training, a pre-training model based on contrasting text-image pairs) to better focus on the application scenarios of image-text relevance in presentations, such as AIPPT (Artificial Intelligence Power Point, software for creating slides and presentations based on artificial intelligence). Fine-tuning can include two stages:
[0116] The first stage, such as Figure 5 As shown in the figure, the image encoder is frozen, and only the text encoder is fine-tuned. Based on the training samples, fine-tuning the text encoder makes it easier to align the features obtained by the image encoder, facilitating faster convergence during the second stage of full fine-tuning.
[0117] In the second stage, the training samples are used to fine-tune the full model parameters of the retrieval model. In order to better extract image features, the resolution of the images in the training samples can be improved in the second stage so that more image features can be extracted.
[0118] During the training phase, model parameters can be adjusted through contrast loss.
[0119] After two stages of fine-tuning, the image and text features are aligned, where the output features of the text and image vectors can be mapped to a unified dimension, such as 512 dimensions.
[0120] In the inference phase, descriptive features of the same length can be obtained by merging the features of the two encoders.
[0121] After the retrieval model is trained, the screening of candidate image sets and illustrations can be completed in the inference stage.
[0122] (1) Screening candidate image sets
[0123] In the embodiment of the present disclosure, the operation for each description statement is the same. For each description statement, the candidate image of the description statement can be obtained by the following method: Figure 6 As shown:
[0124] S601: extract text features of the road description sentence through a text encoder.
[0125] S602: Determine the similarity between the text feature and each image feature in the image feature set.
[0126] S603: Based on the similarity, select an image that matches the text feature as a candidate image.
[0127] The images with higher similarity are selected as candidate images. In implementation, the first m images can be selected as candidate images, or images with similarity higher than a similarity threshold can be selected as candidate images.
[0128] In the disclosed embodiment, the feature extraction capability based on the neural network model can extract key features from the description sentence to facilitate the recall of suitable candidate images, which can help improve the accuracy of the matching images.
[0129] In the embodiment of the present disclosure, when pictures are required, a retrieval model can be used to extract the image features of each image and compare them with the text features of each description sentence.
[0130] In addition, to improve the efficiency of image matching, image feature sets are pre-generated and stored. The image features can be extracted and stored in the cache in advance. This can reduce the need for repeated image feature extraction when matching is needed, thus speeding up the matching process.
[0131] Furthermore, in the embodiment of the present disclosure, in order to meet the needs of matching pictures, the image features stored in the image feature set include not only features extracted by the image encoder, but also text features extracted based on the text of the corresponding image.
[0132] In order to better describe the feature expression of the image, in the embodiment of the present disclosure, the image features of each image to be matched in the set of images to be matched are obtained based on the following method:
[0133] Step A1, extracting a first feature expression of the image to be matched based on an image encoder; and
[0134] Step A2: extracting a second feature expression of the image text corresponding to the image to be matched based on the text encoder.
[0135] It should be noted that the execution order of step A1 and step A2 is not limited.
[0136] Step A3: Perform weighted summation on the first feature expression and the second feature expression to obtain image features of the image to be matched.
[0137] During implementation, the weights can be obtained through training when training the retrieval model. For example, after the second stage of training, the model parameters of the image encoder and the text encoder are frozen, and appropriate weights are trained using training samples. When training the weights, a supervised training method is used to compare the comprehensive features of the candidate text and the image (i.e., the features after weighted summation), determine the similarity between the two, and adjust the weights based on the contrast loss. That is, the weights are trained by maximizing the similarity between the candidate text and image features that express the same or similar content, and minimizing the similarity between the candidate text and image features that express different or dissimilar content. The disclosed embodiments are not limited to this.
[0138] In the disclosed embodiments, the image features used for retrieval combine features extracted by the image encoder and the text encoder, enabling the description of images from multiple dimensions. In particular, when matching text with images, the features described by the text encoder dimensions can better adapt to the text on the page being matched, thereby improving the efficiency of matching images.
[0139] (2) Screening illustrations
[0140] In the embodiment of the present disclosure, since each description statement recalls candidate images separately, there may be duplicate images between the candidate images recalled by each description statement. In order to improve the efficiency of image matching, in the embodiment of the present disclosure, duplicate images are removed from the candidate image set based on the following method:
[0141] Step B1: Obtain image features of each candidate image in the candidate image set.
[0142] Step B2: normalize the image features of each candidate image to obtain the normalized features of each candidate image.
[0143] Through normalization operations, the features of different images can be expressed in a smaller feature space, which can appropriately reduce the amount of calculation and improve the efficiency of deduplication.
[0144] Step B3: construct a feature matrix based on the normalized features of each candidate image.
[0145] For example, in the feature matrix, elements in the same row are associated with each candidate image, and elements in the same column are also associated with each candidate image. Therefore, for different candidate characteristics, the feature matrix can be used to sort out pairs of images so that in step B4, multiple threads created based on at least one core can perform element-by-element multiplication operations on the feature matrix to obtain the similarity between any two candidate images.
[0146] Step B5: The two candidate images with a similarity higher than a preset threshold are regarded as duplicate images, and a duplicate image label is obtained.
[0147] Step B6: performing image deduplication operation on the candidate image set based on the duplicate image tags.
[0148] For duplicate images, only one image is retained in the candidate image set. In practice, the decision on which candidate image to retain can be based on the similarity between the candidate image and the page to be matched. That is, the candidate image with the highest similarity is retained. Alternatively, if duplicate images belong to the same category, the class center vector of that category is found, and the candidate image closest to the class center vector is retained.
[0149] In the disclosed embodiments, multi-channel recall improves the recall rate, but this may also result in a large number of recalled images. To this end, deduplication is used to optimize images and improve image matching efficiency. To minimize the number of images subsequently processed, normalization and kernel multithreading techniques are used to further improve deduplication efficiency, ultimately achieving the goal of improving image matching efficiency.
[0150] In the embodiment of the present disclosure, determining the matching degree between each candidate image in the candidate image set and the page to be matched can be implemented as follows: Figure 7 As shown:
[0151] S701: extract at least one second-category keyword from the text of the page to be matched with an image based on a second keyword extraction method.
[0152] In the embodiment of the present disclosure, the second keyword extraction method is different from the first keyword extraction method. In this way, different keywords are used to describe the illustration requirements at different stages, which can help improve the quality of illustrations.
[0153] For example, the first keyword extraction method may extract keywords based on word frequency, and the second keyword extraction method may use a large model to extract keywords.
[0154] The prompt example of a large model is as follows:
[0155] "Input: #hh and gg ## Comparison of hh and gg's writing styles ### Select similarities and differences -hh's works often deal with social realities, explorations of human nature, and youthful growth. They focus on the lives and fates of those at the bottom of society, conveying profound reflections on society. -gg's works often revolve around themes such as love, friendship, and dreams, focusing on the individual's inner emotional world and growth process, expressing a yearning for and pursuit of beautiful things. The above is text from a PPT page. # represents different levels of titles. \n represents a carriage return. Requires extracting key information from the text to search for relevant images and add them to the PPT. \n\nOutput: \n Add a photo of hh / gg. \n\n\nInput: \n{}\n\nThe above is text from a PPT page. # represents different levels of titles. \n represents a carriage return. Requires extracting key information from the text to search for relevant images and add them to the PPT. \n\nOutput: \n"
[0156] That is, the prompt gives an example of a page that needs to be illustrated, and explains the title and task in the example. Based on this example, the large model can be guided to extract keywords for the page to be illustrated.
[0157] S702: Splice the second category of keywords into the text of the page to be matched with pictures to obtain the intermediate text.
[0158] S703: Use a text encoder to extract text features of the intermediate text.
[0159] S704 , calculating the matching degree between the text features of the intermediate text and the image features of each candidate image in the candidate image set, and obtaining the matching degree between each candidate image in the candidate image set and the page to be matched.
[0160] Afterwards, suitable illustrations can be filtered out based on the matching degree to complete the illustration operation.
[0161] In the disclosed embodiment, further expanded keywords are used to assist in improving the text description of the page to be matched with pictures, which can help the text encoder further understand the content of the page to be matched with pictures, extract reasonable text features, and thus find suitable illustrations.
[0162] In the embodiments of the present disclosure, the experimental results can be further described through comparative experiments. Figure 8 As shown, the left column of pictures is the effect of the illustration without using the scheme provided by the embodiment of the present disclosure, and the right column is the effect of the illustration using the scheme provided by the embodiment of the present disclosure. Figure 8 From the comparison diagram in the first row, it can be seen that the diagrams in the embodiment of the present disclosure can intuitively express the scene to be described. Figure 8 As can be seen from the comparison chart in the second row, the illustration of the logistics scene in the left picture is relatively simple, while the illustration provided by the embodiment of the present disclosure has a richer display of the logistics scene and can better fit the text language.
[0163] Based on the same technical concept, the embodiment of the present disclosure also provides an automatic presentation diagram device 900, such as Figure 9 Shown, including:
[0164] Extraction module 901, used to extract multiple description sentences from the page to be matched with pictures in the presentation;
[0165] A matching module 902 is configured to match each description statement with a corresponding candidate image to obtain a candidate image set;
[0166] Determination module 903, used to determine the matching degree between each candidate image in the candidate image set and the page to be matched;
[0167] The screening module 904 is used to screen out target images for the page to be matched based on the matching degree between each candidate image and the page to be matched, so as to serve as illustrations for the page to be matched.
[0168] In some embodiments, the extraction module includes:
[0169] An extraction unit, used to extract the title information and text content of the page to be matched with the image from the presentation;
[0170] A generating unit is used to generate the multi-path description sentence based on the title information and the text content.
[0171] In some embodiments, the generating unit is configured to perform at least two of the following operations to generate the multi-path description statement:
[0172] According to the level of the title, the titles of each level in the title information are sequentially spliced to obtain the first description statement;
[0173] The first description statement is concatenated with the main text content in a preset order to obtain a second description statement;
[0174] Based on the first keyword extraction method, extract at least one first-category keyword from the title information and the body content to construct a third description statement;
[0175] Target-type words are extracted in order from low to high levels of the titles; and a fourth description statement is constructed based on the extracted target-type words; wherein, if the target-type words are extracted from a title of any level, the target-type words will no longer be extracted from titles of a level higher than that of the title of any level.
[0176] In some embodiments, the method further includes adding a module, including:
[0177] Extracting a target proper name from a preset proper name type set from the title information and the text content;
[0178] When the target proper name is extracted, the target proper name is added to at least one description statement in the multi-way description statement.
[0179] In some embodiments, the generating unit is specifically configured to:
[0180] An explanation text in a preset format is added to the target word to obtain the fourth description sentence.
[0181] In some embodiments, the matching module is specifically configured to:
[0182] For each description sentence, extract the text features of the description sentence through a text encoder;
[0183] Determine the similarity between the text feature and each image feature in the image feature set;
[0184] Based on the similarity, images matching the text feature are selected as candidate images.
[0185] In some embodiments, an image feature extraction module is further included, which is used to:
[0186] Extracting a first feature expression of the image to be matched based on an image encoder; and
[0187] Extracting a second feature expression of the image text corresponding to the image to be matched based on the text encoder;
[0188] Performing a weighted summation on the first feature expression and the second feature expression to obtain the image feature of the image to be matched.
[0189] In some embodiments, the determination module is specifically configured to:
[0190] Extracting at least one second-category keyword from the text of the page to be matched with an image based on the second keyword extraction method;
[0191] Splicing the second category of keywords into the text of the page to be matched with pictures to obtain an intermediate text;
[0192] Using a text encoder to extract text features of the intermediate text;
[0193] The matching degree of the text features of the intermediate text and the image features of each candidate image in the candidate image set is calculated respectively, and the matching degree of each candidate image in the candidate image set and the page to be matched with an image is obtained.
[0194] In some embodiments, a deduplication module is further included for:
[0195] Obtaining image features of each candidate image in the candidate image set;
[0196] Normalizing the image features of each candidate image to obtain normalized features of each candidate image;
[0197] Based on the normalized features of each candidate image, a feature matrix is constructed;
[0198] Multiple threads created based on at least one core perform product operations between elements of the feature matrix to obtain the similarity between any two candidate images;
[0199] The two candidate images whose similarity is higher than a preset threshold are regarded as duplicate images and a duplicate image mark is obtained;
[0200] An image deduplication operation is performed on the candidate image set based on the duplicate image tags.
[0201] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0202] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0203] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0204] Figure 10A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0205] like Figure 10 As shown, the device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0206] Various components in device 1000 are connected to I / O interface 1005, including an input unit 1006, such as a keyboard, mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, optical disk, etc.; and a communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0207] The computing unit 1001 can be various general and / or special purpose processing components having processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs various methods and processes described above, such as the automatic graphing method of a presentation. For example, in some embodiments, the automatic graphing method of a presentation can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded onto the RAM 1003 and executed by the computing unit 1001, one or more steps of the automatic graphing method of a presentation described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the automatic graphing method of a presentation by any other suitable means, such as by means of firmware.
[0208] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0209] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a function / operation specified in the flowchart and / or block diagram. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, and partially on a remote machine or a server.
[0210] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0211] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0212] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0213] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0214] Based on the aforementioned electronic devices, the present disclosure also provides a vehicle, which may include electronic devices, and may also include communication components, a display screen for realizing a human-machine interface, and an information collection device for collecting surrounding environment information, etc. The communication components, the display screen, the information collection device and the electronic devices are communicatively connected.
[0215] According to an embodiment of the present disclosure, the electronic device may be integrated with the communication component, the display screen, and the information collection device, or may be separately provided with the communication component, the display screen, and the information collection device.
[0216] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0217] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for automatically adding pictures to a presentation, comprising: Extract multiple description sentences from the page to be matched with pictures in the presentation; Match each description statement with the corresponding candidate image to obtain a candidate image set; Determining the matching degree between each candidate image in the candidate image set and the page to be matched with an image; Based on the matching degree between each candidate image and the page to be illustrated, a target image is selected for the page to be illustrated as an illustration of the page to be illustrated; The step of extracting multiple description sentences from the page of the presentation to be provided with pictures includes performing at least two of the following operations: According to the level of the title, the titles of each level in the title information are sequentially spliced to obtain the first description statement; splicing the first description statement with the main text content in a preset order to obtain a second description statement; Based on the first keyword extraction method, extract at least one first-category keyword from the title information and the body content to construct a third description statement; Extract target-class words in order of the titles' levels from low to high; and based on the extracted target-class words, add pre-formatted explanatory text to the target-class words to construct a fourth description statement; wherein, if the target-class words are extracted from any level of title, the target-class words are no longer extracted from titles with a level higher than that of the title; The explanatory text is defined based on a template, or is generated based on information generated by processing the page to be illustrated based on a natural language understanding model.
2. The method according to claim 1, wherein The step of extracting multiple description sentences from the page of the presentation to be equipped with pictures includes: Extract the title information and text content of the page to be matched with the image from the presentation; The multi-path description sentence is generated based on the title information and the body content.
3. The method according to claim 1 or 2, further comprising: Extracting a target proper noun from a preset proper noun type set from the title information and the text content; In the case where the target proper name is extracted, the target proper name is added to at least one description statement in the multiple description statements.
4. The method according to claim 1, wherein The step of matching each description statement to a corresponding candidate image includes: For each description sentence, extract the text features of the description sentence through a text encoder; Determining the similarity between the text feature and each image feature in the image feature set; Based on the similarity, an image matching the text feature is selected as a candidate image.
5. The method according to claim 4, wherein The method also includes obtaining the image features of each image to be matched in the image set to be matched based on the following method: extracting a first feature expression of the image to be matched based on an image encoder; as well as, Extracting a second feature expression of the image text corresponding to the image to be matched based on the text encoder; Perform a weighted summation on the first feature expression and the second feature expression to obtain the image feature of the image to be matched.
6. The method according to claim 1, wherein Determining the matching degree between each candidate image in the candidate image set and the page to be matched includes: Extracting at least one second-category keyword from the text of the page to be matched with an image based on a second keyword extraction method; Splicing the second category of keywords into the text of the page to be matched with pictures to obtain an intermediate text; extracting text features of the intermediate text using a text encoder; The text features of the intermediate text and the image features of each candidate image in the candidate image set are respectively calculated to obtain the matching degree of each candidate image in the candidate image set with the page to be matched with an image.
7. The method according to claim 1, further comprising: Obtaining image features of each candidate image in the candidate image set; Normalizing the image features of each candidate image to obtain normalized features of each candidate image; Based on the normalized features of each candidate image, a feature matrix is constructed; performing a product operation between elements of the feature matrix based on multiple threads created by at least one core to obtain a similarity between any two candidate images; The two candidate images whose similarity is higher than a preset threshold are regarded as duplicate images and a duplicate image mark is obtained; An image deduplication operation is performed on the candidate image set based on the duplicate image tags.
8. An automatic diagram-matching device for presentations, comprising: An extraction module is used to extract multiple description sentences from the page to be matched with pictures in the presentation; A matching module is used to match the corresponding candidate images for each description statement to obtain a candidate image set; A determination module, configured to determine a degree of matching between each candidate image in the candidate image set and the page to be matched with an image; A screening module, configured to screen out a target image for the page to be illustrated based on the degree of matching between each candidate image and the page to be illustrated; The step of extracting multiple description sentences from the page of the presentation to be provided with pictures includes performing at least two of the following operations: According to the level of the title, the titles of each level in the title information are sequentially spliced to obtain the first description statement; splicing the first description statement with the main text content in a preset order to obtain a second description statement; Based on the first keyword extraction method, extract at least one first-category keyword from the title information and the body content to construct a third description statement; Extract target-class words in order of the titles' levels from low to high; and based on the extracted target-class words, add pre-formatted explanatory text to the target-class words to construct a fourth description statement; wherein, if the target-class words are extracted from any level of title, the target-class words are no longer extracted from titles with a level higher than that of the title; The explanatory text is defined based on a template, or is generated based on information generated by processing the page to be illustrated based on a natural language understanding model.
9. The device according to claim 8, wherein The extraction module comprises: An extraction unit, used to extract the title information and text content of the page to be matched with the image from the presentation; A generating unit is configured to generate the multi-path description sentences based on the title information and the body content.
10. The apparatus according to claim 8 or 9, further comprising an adding module, comprising: Extracting a target proper noun from a preset proper noun type set from the title information and the text content; In the case where the target proper name is extracted, the target proper name is added to at least one description statement in the multiple description statements.
11. The device according to claim 8, wherein The matching module is specifically used for: For each description sentence, extract the text features of the description sentence through a text encoder; Determining the similarity between the text feature and each image feature in the image feature set; Based on the similarity, an image matching the text feature is selected as a candidate image.
12. The device according to claim 11, wherein It also includes an image feature extraction module for: extracting a first feature expression of the image to be matched based on the image encoder; as well as, Extracting a second feature expression of the image text corresponding to the image to be matched based on the text encoder; Perform a weighted summation on the first feature expression and the second feature expression to obtain the image feature of the image to be matched.
13. The device according to claim 8, wherein The determining module is specifically configured to: Extracting at least one second-category keyword from the text of the page to be matched with an image based on a second keyword extraction method; Splicing the second category of keywords into the text of the page to be matched with pictures to obtain an intermediate text; extracting text features of the intermediate text using a text encoder; The text features of the intermediate text and the image features of each candidate image in the candidate image set are respectively calculated to obtain the matching degree of each candidate image in the candidate image set with the page to be matched with an image.
14. The apparatus according to claim 8, further comprising a deduplication module, configured to: Obtaining image features of each candidate image in the candidate image set; Normalizing the image features of each candidate image to obtain normalized features of each candidate image; Based on the normalized features of each candidate image, a feature matrix is constructed; performing a product operation between elements of the feature matrix based on multiple threads created by at least one core to obtain a similarity between any two candidate images; The two candidate images whose similarity is higher than a preset threshold are regarded as duplicate images and a duplicate image mark is obtained; An image deduplication operation is performed on the candidate image set based on the duplicate image tags.
15. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
17. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
News illustration method and device, equipment and storage medium
CN113343012A
Method for generating inforgraphic information and method for generating image database
WO2020103899A1
Presentation generation method and apparatus, computer device and storage medium
WO2021164255A1