Media preview method and apparatus, computer device, and storage medium
Patent Information
- Application Number
- CN202211018035.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-24
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-08-24
AI Technical Summary
然而,对于包含大量文字的图片,如公告、文字截图等,经过压缩后的缩略图中文字往往难以辨认,导致用户在预览时难以区别各图片,需要点击查看原图后才能够准确进行区分,媒体预览时所呈现的信息有限
[0024]上述媒体预览方法、装置、计算机设备、存储介质和计算机程序产品,对于文本显示区域占比达到文本主题占比阈值的文本类型图像,在预览的图像缩略图中,显示将截取自文本类型图像的部分区域缩放到预设尺寸后的缩放图像,缩放图像中文本的分辨率不小于文本可视分辨率阈值,且部分区域内的文本重要程度相比部分区域外的文本重要程度更高,从而能够通过预览的缩放图像直接展示出文本重要程度高的文本内容,增加了媒体预览时所呈现的信息量,有利于用户根据预览的缩略图选择所需查看的媒体资源。
Smart Images

Figure CN117010325B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a media preview method, apparatus, computer device, storage medium, and computer program product. Background Technology
[0002] Currently, the media library, which includes media resources such as videos and images, provides a preview function for users to preview and select the media they wish to view. During media preview, users can see the overall content of the media through compressed thumbnails. However, for images containing a large amount of text, such as announcements or screenshots, the text in the compressed thumbnails is often difficult to discern, making it hard for users to distinguish between images during preview. Users often need to click to view the original image to accurately differentiate between them, thus limiting the information presented during media preview. Summary of the Invention
[0003] Therefore, it is necessary to provide a media preview method, apparatus, computer device, computer-readable storage medium, and computer program product that can increase the amount of information presented during media preview, in order to address the aforementioned technical problems.
[0004] Firstly, this application provides a media preview method. The method includes:
[0005] In response to a preview trigger event for the media library, display the media preview area of the media library;
[0006] When the media library includes text-type images, a thumbnail pointing to the text-type image is displayed in the media preview area; the image thumbnail has a preset size, and the proportion of the text display area in the text-type image reaches the text theme proportion threshold;
[0007] In the image thumbnail pointing to the text-type image, a scaled image is displayed after a portion of the text-type image has been cropped to a preset size; the resolution of the text in the scaled image is not less than the text visibility resolution threshold; in the text-type image, the text within a portion of the image is more important than the text outside that portion.
[0008] Secondly, this application also provides a media preview device. The device includes:
[0009] The preview area display module is used to display the media preview area of the media library in response to preview trigger events for the media library;
[0010] The thumbnail display module is used to display image thumbnails pointing to text-type images in the media preview area when the media library includes text-type images; the image thumbnails have a preset size, and the proportion of the text display area in the text-type image reaches the text theme proportion threshold;
[0011] The text region display module is used to display a scaled image of a portion of the text-type image, which is then scaled to a preset size, in an image thumbnail pointing to a text-type image; the resolution of the text in the scaled image is not less than the text visibility resolution threshold; and in the text-type image, the text within a portion of the image is more important than the text outside that portion.
[0012] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to perform the following steps:
[0013] In response to a preview trigger event for the media library, display the media preview area of the media library;
[0014] When the media library includes text-type images, a thumbnail pointing to the text-type image is displayed in the media preview area; the image thumbnail has a preset size, and the proportion of the text display area in the text-type image reaches the text theme proportion threshold;
[0015] In the image thumbnail pointing to the text-type image, a scaled image is displayed after a portion of the text-type image has been cropped to a preset size; the resolution of the text in the scaled image is not less than the text visibility resolution threshold; in the text-type image, the text within a portion of the image is more important than the text outside that portion.
[0016] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, performs the following steps:
[0017] In response to a preview trigger event for the media library, display the media preview area of the media library;
[0018] When the media library includes text-type images, a thumbnail pointing to the text-type image is displayed in the media preview area; the image thumbnail has a preset size, and the proportion of the text display area in the text-type image reaches the text theme proportion threshold;
[0019] In the image thumbnail pointing to the text-type image, a scaled image is displayed after a portion of the text-type image has been cropped to a preset size; the resolution of the text in the scaled image is not less than the text visibility resolution threshold; in the text-type image, the text within a portion of the image is more important than the text outside that portion.
[0020] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, performs the following steps:
[0021] In response to a preview trigger event for the media library, display the media preview area of the media library;
[0022] When the media library includes text-type images, a thumbnail pointing to the text-type image is displayed in the media preview area; the image thumbnail has a preset size, and the proportion of the text display area in the text-type image reaches the text theme proportion threshold;
[0023] In the image thumbnail pointing to the text-type image, a scaled image is displayed after a portion of the text-type image has been cropped to a preset size; the resolution of the text in the scaled image is not less than the text visibility resolution threshold; in the text-type image, the text within a portion of the image is more important than the text outside that portion.
[0024] The aforementioned media preview method, apparatus, computer equipment, storage medium, and computer program products, for text-type images where the text display area occupies a threshold value for the text theme, display a scaled image in the preview image thumbnail after scaling a portion of the text-type image to a preset size. The resolution of the text in the scaled image is not less than the text visual resolution threshold, and the text within the portion of the image is more important than the text outside the portion of the image. This allows for the direct display of text content with high text importance through the scaled preview image, increasing the amount of information presented during media preview and facilitating users in selecting the media resources they wish to view based on the preview thumbnail. Attached Figure Description
[0025] Figure 1 This is a diagram illustrating the application environment of the media preview method in one embodiment;
[0026] Figure 2 This is a flowchart illustrating a media preview method in one embodiment;
[0027] Figure 3 This is a schematic diagram of an interface displaying image thumbnails in a media preview interface of one embodiment;
[0028] Figure 4 This is a schematic diagram of an interface displaying keywords in an image thumbnail in one embodiment;
[0029] Figure 5 This is a flowchart illustrating the text type image determination process in one embodiment;
[0030] Figure 6 This is a schematic diagram of the interface for browsing the photo album in one embodiment;
[0031] Figure 7 This is a schematic diagram of the interface for browsing media in an application, as shown in one embodiment.
[0032] Figure 8 This is a schematic diagram of a text-type image in one embodiment;
[0033] Figure 9 for Figure 8 The image thumbnails corresponding to the text type images in the illustrated embodiments;
[0034] Figure 10 This is a schematic diagram of a screenshot from one embodiment;
[0035] Figure 11 for Figure 10 The image thumbnails corresponding to the article screenshots in the illustrated embodiments;
[0036] Figure 12 This is a schematic diagram of an interface displaying text keywords associated with a text-type image in one embodiment;
[0037] Figure 13 This is a flowchart illustrating the media preview method in another embodiment;
[0038] Figure 14 for Figure 10 The illustrated embodiment is a diagram showing the text region identified from a screenshot of an article.
[0039] Figure 15 This is a schematic diagram illustrating center-cropping based on height priority in one embodiment;
[0040] Figure 16 This is a schematic diagram of a width-priority center-cropping method in one embodiment;
[0041] Figure 17 This is a schematic diagram illustrating the centering of a key area based on width priority in one embodiment;
[0042] Figure 18 This is a structural block diagram of a media preview device in one embodiment;
[0043] Figure 19 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0045] The media preview method provided in this application embodiment can be applied to, for example, Figure 1In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another server. The media library can be a local media resource library of terminal 102, or a media resource library of server 104 accessed by terminal 102 via a network. The media library can include media resources obtained locally by terminal 102, such as locally captured images or videos, and can also include media resources obtained by terminal 102 from server 104 via a network. Users can trigger a preview event for the media library on terminal 102. Terminal 102 displays a media preview area of the media library. For text-type images that meet the text theme conditions (i.e., for text-type images where the text display area reaches the text theme proportion threshold), terminal 102 displays a scaled image in the preview image thumbnail, where a portion of the text-type image is scaled to a preset size. The resolution of the text in the scaled image is not less than the text visual resolution threshold, and the text within the scaled image is more important than the text outside the scaled image.
[0046] The terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle systems. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0047] In one embodiment, such as Figure 2 As shown, a media preview method is provided. This method is executed by a computer device, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking the terminal in the example, the explanation includes the following steps:
[0048] Step 202: In response to a preview trigger event for the media library, display the media preview area of the media library.
[0049] Media can include various resources such as images and videos. A media library is a resource database of media, containing various media that can be accessed and viewed. Previewing refers to an interactive method of viewing media in advance without directly accessing it. A preview trigger event is an event that triggers a preview of the media in the media library. Preview trigger events can be generated by user actions, such as a user triggering a preview operation on the media library, thus generating a preview trigger event. Preview trigger events can also be generated automatically when preset trigger conditions are met, such as reaching a preset time, a preset location, or a preset number of media, automatically generating a preview trigger event to preview the media library. The media preview area is an area for previewing each media in the media library. The media preview area can display thumbnails of each media in the media library, allowing users to preview the media in the library by browsing these thumbnails.
[0050] Specifically, users can trigger operations on the media library of a terminal or server to generate preview trigger events, indicating that they need to preview each media in the media library. The terminal responds to the preview trigger event and displays the media preview area of the media library so that the user can preview each media in the media preview area.
[0051] Step 204: When the media library includes text-type images, display an image thumbnail pointing to the text-type image in the media preview area; the image thumbnail has a preset size, and the proportion of the text display area in the text-type image reaches the text theme proportion threshold.
[0052] The text display area ratio refers to the proportion of text within a text-type image, specifically calculated based on the relative sizes of the text display area and the text-type image area. The text theme ratio threshold determines whether an image's theme is text content. This threshold can be set according to specific needs, such as different thresholds for images from different sources. Furthermore, if the text display area ratio in a text-type image reaches the text theme ratio threshold, it means the relationship between the text in the image and the text-type image meets the text theme condition; the text-type image has text content as its theme, and the theme of the image is its included text content. For example, a screenshot of an article on the internet, while its media format is an image, if the text captured in the image is the theme of the image, then the relationship between the text in the screenshot and the screenshot meets the text theme condition. Similarly, a text-type image, while its media format is an image, has its theme content as the text within it; that is, the text in the image is the core information carried by the image. In practical applications, images in a media library can be analyzed to determine whether they belong to the text-type image category, such as whether the relationship between the text in the image and the image meets the text theme condition. In a practical implementation, the proportion of the text display area in the image can be compared with a preset text theme proportion threshold to determine whether the text theme condition is met, that is, whether the image has the included text as its theme, or whether the image belongs to the text type image.
[0053] Image thumbnails have preset sizes, which can be flexibly set according to actual needs. For example, they can be 1 M pixels wide and 1 M pixels wide. Each image thumbnail points to a media item in the media library, displaying a preview of that media. Compared to the original size of the media, the image thumbnail is scaled down to a suitable size for previewing. Users can trigger actions on the image thumbnails, such as clicking on them to access the corresponding media.
[0054] Specifically, in the media preview area of the media library, the terminal displays image thumbnails pointing to each media item. These thumbnails have preset sizes, which can be set according to actual needs. Users can also customize the preset sizes of the image thumbnails to meet specific preview requirements. For text-type images included in the media library, the terminal displays an image thumbnail pointing to that text-type image in the media preview area. The relationship between the text in the text-type image and the text-type image conforms to the text theme condition, that is, the text in the text-type image serves as the main content of that text-type image.
[0055] Step 206: In the image thumbnail pointing to the text type image, display a scaled image after scaling a portion of the text type image to a preset size; the resolution of the text in the scaled image is not less than the text visibility resolution threshold; in the text type image, the text within a portion of the image is more important than the text outside that portion of the image.
[0056] The image includes a region containing text, where the text within this region is more important than the text outside this region. Text importance is used to characterize the significance of the text within the text-type image. In a text-type image, regions with higher text importance are considered more important overall. Since the text within a region is more important than the text outside this region, this region can be the area containing the most important text in the text-type image. The text resolution in the scaled image must be no less than the visible resolution threshold to meet the text recognition criteria. This means that when the scaled image is displayed as a thumbnail, the text in the scaled image should be recognizable to the user, allowing them to accurately identify the text content. The text recognition criteria can be set according to actual needs, such as setting the font size to meet a font size threshold. The visible resolution threshold can be the minimum resolution required to accurately recognize the text content. The text resolution can be determined based on the font size, such as directly using the font height as the text resolution. If the font height in the scaled image is greater than or equal to the visible resolution threshold, it ensures that the text in the scaled image can be accurately recognized by the user. The scaled image is the same size as the image thumbnail. The scaled image is obtained by scaling a portion of the text-type image according to the preset size of the image thumbnail. The scaled image is used to represent the corresponding text-type image.
[0057] Specifically, in the image thumbnail, the terminal displays a scaled-down image. The resolution of the text in the scaled-down image is no less than the text visual resolution threshold, meaning the text in the scaled-down image can be accurately recognized by the user. The scaled-down image is obtained by cropping a portion of the text-type image and scaling it to a preset size. This portion represents the area with the highest text importance in the text-type image; that is, the text within this portion is more important than the text outside this portion. Through the text in this portion, the main content of the text-type image can be accurately represented.
[0058] For example, for a text-type image, specifically a screenshot of a complete article, the area cropped from the text-type image can include the title. After scaling the title-included area to a preset size, the resulting scaled image, with a resolution no less than the text visibility resolution threshold, is displayed in the image thumbnail. Users can directly identify the article title in the text-type image based on the thumbnail during preview, thus obtaining the article's main content. This allows users to select the desired media resource based on the preview thumbnail, avoiding the need to view the original text-type image first, simplifying media selection and improving the user experience.
[0059] In a specific application, such as Figure 3 As shown, the media preview interface in the media library displays image thumbnails pointing to media within the library. These thumbnails can point to images or videos within the library. Images 1, 2, and 4 are text-type images, and the relationship between the text and the text in these images conforms to the text theme condition; that is, images 1, 2, and 4 all have text content as their theme. Image 3 is a photograph of a person. Each image thumbnail displayed in the media preview interface points to its corresponding media. For images 1, 2, and 4, which are text-type images, scaled images are displayed, cropped from a portion of each image and scaled to a preset size. The resolution of the text in the scaled images is not less than the text visual resolution threshold, and the text within certain areas is more important than the text outside those areas. In the image thumbnails... Figure 1 The image displays the title of the article in the screenshot, allowing users to quickly grasp the key information in Image 1, namely, content related to daily health and wellness. (Image thumbnail) Figure 2 and image thumbnails Figure 4 In the media preview interface, each image displays the content of the group announcement. Users can directly access the key information in images 2 and 4, which pertains to the content of the group announcement.
[0060] In the aforementioned media preview method, for text-type images where the text display area occupies a threshold value for the text theme, a scaled image is displayed in the preview image thumbnail after a portion of the text-type image has been scaled up to a preset size. The resolution of the text in the scaled image is not less than the text visual resolution threshold, and the text within a portion of the image is more important than the text outside that portion. This allows for the direct display of text content with high importance through the scaled preview image, increasing the amount of information presented during media preview and facilitating users in selecting the media resources they wish to view based on the preview thumbnail.
[0061] In one embodiment, the media preview method further includes: displaying at least one text keyword associated with the text-type image in the image thumbnail; the at least one text keyword is used to describe the content theme of the text in the text-type image.
[0062] The text keywords are associated with text-type images. Specifically, the text keywords can be obtained based on text recognition within the text-type image and are used to describe the content theme of the text in the image. Text keywords can be extracted directly from the text-type image or obtained through semantic recognition of the text within the image. There must be at least one text keyword, and the number can be set according to actual needs. Specifically, it can be set to a fixed number, such as 3; or it can be set to a range, such as 3-5, in which case 3-5 corresponding text keywords will be displayed based on the text keywords recognized from the text in the text-type image.
[0063] Specifically, the terminal displays at least one text keyword associated with the text-type image in the image thumbnail. The text keyword can be displayed overlaid on the scaled image. The resolution of the text keyword is not less than the text visual resolution threshold; that is, the at least one text keyword displayed in the image thumbnail should be easily identifiable by the user during preview, allowing for accurate differentiation of each text-type image based on its text keywords. The text keyword describes the content theme of the text in the corresponding text-type image. Different texts in different text-type images have different content themes, resulting in different associated text keywords. The text keyword displays key information about the text-type image, enabling users to preview and select from various thumbnails based on this key information.
[0064] In a specific application, such as Figure 4 As shown, the media preview interface in the media library displays four image thumbnails from left to right. For the first and second image thumbnails, text keywords are displayed for each image. The text keywords for the first image include "lifestyle, health, exercise, diet," while the text keyword for the second image includes "group announcement." Users can quickly understand the important content information of the image linked to by viewing the text keywords associated with the corresponding image in the thumbnail. The third image thumbnail, a photograph of a person, is not a text-based image and is displayed directly as a scaled-down image in the media preview interface.
[0065] In this embodiment, the terminal displays at least one text keyword describing the content theme of the text in the image thumbnail, enabling users to accurately identify the content based on the text keyword during preview. This further increases the amount of information presented during media preview by using the text keyword, which is beneficial for users to select the media resources they want to view based on the preview thumbnail.
[0066] In one embodiment, displaying at least one text keyword associated with a text-type image in an image thumbnail includes: in response to a keyword display triggering operation in a media preview area, displaying at least one text keyword associated with a text-type image from which the zoomed image originates in a zoomed image.
[0067] The keyword display trigger operation is the operation that triggers the display of text keywords associated with a text-type image. Keyword display triggers can be user-initiated, such as by triggering an operation on a keyword display control in the media library. The text keywords associated with the text-type image are displayed in the scaled image, specifically overlaid on top of the scaled image, or floating above it.
[0068] Specifically, users can trigger a keyword display operation in the media preview area. For example, if a user clicks the keyword display control in the media preview area, the terminal responds to the user's keyword display trigger operation and displays at least one text keyword in the zoomed image. The displayed text keyword is associated with the text type image from which the zoomed image originates, and can be specifically identified from the text in the text type image from which the zoomed image originates.
[0069] In this embodiment, the terminal responds to the user's keyword display trigger operation in the media preview area and displays at least one text keyword associated with the text type image in the zoomed image. The displayed text keyword describes the text type image pointed to by the image thumbnail. According to the user's needs, the amount of information presented during media preview can be increased by using text keywords, which is beneficial for the user to select the media resources to be viewed based on the preview thumbnail.
[0070] In one embodiment, the media preview method further includes: displaying an activation entry point in the media preview area by showing keywords that indicate the activation status;
[0071] The keyword display activation entry is the user's entry point to trigger the display of text keywords. This entry also indicates the activation status of the displayed text keywords, which can be displayed using different methods such as different colors or markers. When the keyword display activation entry is in a pending activation state, it means the text keyword is not currently activated, i.e., the text keyword associated with the text-type image is not displayed in the image thumbnail. When the keyword display activation entry is in an activated state, it means the text keyword is currently activated, i.e., the text keyword associated with the text-type image is already displayed in the image thumbnail. Through the keyword display activation entry, users can control whether or not to display text keywords associated with text-type images, and the status of the entry indicates whether the text keyword associated with the text-type image is currently activated.
[0072] Specifically, the terminal can display a keyword display activation entry in the media preview area. The specific form of this keyword display activation entry can be set according to actual needs, such as a switch control, slider control, etc. The keyword display activation entry can indicate the current status of the keyword display, specifically indicating the pending activation status, to show that the text keyword associated with the text-type image has not yet been activated.
[0073] Furthermore, in response to a keyword display triggering operation in the media preview area, at least one text keyword associated with the text type image from which the zoomed image originates is displayed in the zoomed image, including: in response to a triggering operation on the keyword display activation entry, switching the keyword display activation entry to an active state, and displaying at least one text keyword associated with the text type image from which the zoomed image originates in the zoomed image.
[0074] The triggering action is the user's action that activates the entry point for displaying the keyword, such as a user clicking on the entry point. The activation status indicator identifies the currently activated text keyword associated with the displayed text-type image.
[0075] Specifically, users can trigger an action by displaying an activation entry for a keyword. For example, if a user clicks on the keyword activation entry, the terminal responds by switching the display of the keyword activation entry to indicate its activation status, thus identifying the currently activated text keyword associated with the text-type image. The terminal displays at least one text keyword associated with the text-type image in a zoomed-in image. The text-type image is the image from which the corresponding portion of the zoomed-in image originates.
[0076] Furthermore, the media preview method also includes: in response to a triggering operation of displaying an activation entry for a keyword indicating an active state, switching the display of the keyword activation entry to indicate an active state, and hiding at least one text keyword in the zoomed image.
[0077] The keyword display activation entry is switched to an "Active" state, indicating that the cancellation of the text keyword associated with the text-type image has been triggered. Specifically, the user can trigger the keyword display activation entry again, indicating that the user needs to cancel the display of the text keyword associated with the text-type image. In response to the user's triggering of the keyword display activation entry, the terminal hides at least one text keyword in the zoomed-in image and switches the keyword display activation entry back to an "Active" state. By hiding the displayed text keyword, the terminal displays only the zoomed-in image in the media preview area according to the user's needs.
[0078] In a specific application, such as Figure 4 As shown, a keyword display activation entry is displayed in the media preview area of the media library. This entry is a toggle control, which users can interact with to toggle the activation or deactivation of keywords. The toggle control can indicate whether a keyword is activated or deactivated through different display states. When the toggle control is in the "active" state, users can interact with it by clicking it. The toggle control will then switch to the "active" state, and at least one text keyword associated with the corresponding text type image will be displayed in the zoomed-in image thumbnail. Users can then interact with the toggle control again to deactivate the keyword display. In this case, the toggle control will switch back to the "active" state, and the displayed text keywords will be hidden.
[0079] In this embodiment, the terminal displays a keyword display activation entry in the media preview area. The status of the keyword display activation entry indicates whether the text keywords associated with the text-type images are currently activated. The terminal also supports user-triggered operations on the keyword display activation entry to activate or deactivate the text keywords associated with the text-type images as needed. This allows users to choose whether to increase the amount of information presented during media preview by displaying text keywords, which is beneficial for users to select the media resources they want to view based on the preview thumbnails.
[0080] In one embodiment, displaying at least one text keyword associated with a text-type image in an image thumbnail includes: displaying a preset number of text keywords associated with a text-type image in the image thumbnail; and displaying the preset number of text keywords in order of their respective text importance.
[0081] The preset quantity can be set according to actual needs; it can be a fixed number or a range. For example, a preset quantity of 3 will display 3 text keywords in the image thumbnail. A preset quantity of 3-6 will display 3-6 text keywords in the image thumbnail. Text importance is used to characterize the importance of text in a text-based image. In a text-based image, the higher the text importance, the greater the overall importance of the text. Different regions in a text-based image can have corresponding text importance levels to characterize the importance of different regions within the text-based image; similarly, different text keywords associated with a text-based image can also have corresponding text importance levels to characterize the importance of those keywords within the text-based image.
[0082] Specifically, the terminal displays a preset number of text keywords associated with the text-type image in the image thumbnail, such as 3 or 5 text keywords. Each text keyword can be sorted and displayed according to its own text importance, such as sorting and displaying them in descending order of text importance, thus prioritizing the display of text keywords with higher text importance.
[0083] In this embodiment, the terminal will display a preset number of text keywords in order of their importance. This allows the text keywords to be displayed in an orderly manner based on their importance, thereby increasing the amount of information presented during media preview and helping users select the media resources they want to view based on the preview thumbnails.
[0084] In one embodiment, displaying at least one text keyword associated with a text-type image in an image thumbnail includes: displaying at least one text keyword associated with a text-type image in an image thumbnail when the resolution of the text in the scaled image is less than a text visual resolution threshold.
[0085] Specifically, if the resolution of the text in the scaled image is less than the text visibility resolution threshold, it indicates that the text in the scaled image meets the text recognition criteria. That is, when displayed as a thumbnail, the text in the scaled image can be recognized by the user, meaning the user can still accurately identify the content of the text in the scaled image. The text visibility resolution threshold can be set according to actual needs, such as setting it to a font height of N pixels.
[0086] Specifically, if the resolution of the text in the zoomed image is less than the text visual resolution threshold, it indicates that the user cannot accurately identify the text in the zoomed image when the terminal displays the zoomed image. In this case, the terminal displays at least one text keyword associated with the text type image in the image thumbnail, and the resolution of the displayed text keyword is greater than or equal to the text visual resolution threshold. Thus, even if the user cannot identify the text in the zoomed image, they can still effectively distinguish between the various image thumbnails through the text keyword.
[0087] In a specific application, such as Figure 4 As shown, in the media preview area of the media library, when the resolution of text in a scaled image is less than the text visibility resolution threshold, i.e., the text cannot be accurately recognized by the user, such as in an image thumbnail... Figure 1 and image thumbnails Figure 2 In a zoomed-in image where the user cannot directly discern the exact text content, at least one text keyword associated with the corresponding text type image is displayed floating over the zoomed-in image. For image thumbnails... Figure 4 If the resolution of the text is greater than or equal to the text visual resolution threshold, then the text keywords may not be displayed.
[0088] In this embodiment, when the resolution of the text in the zoomed image is less than the text visual resolution threshold, that is, when the user cannot effectively distinguish between different text types by directly recognizing the text in the zoomed image, the terminal displays at least one text keyword associated with the text type image in the image thumbnail. This enables the user to accurately identify the text based on the text keyword during preview. The text keyword further increases the amount of information presented during media preview, which is beneficial for the user to select the media resource to view based on the preview thumbnail.
[0089] In one embodiment, displaying at least one text keyword associated with a text-type image in an image thumbnail includes: displaying at least one text keyword associated with a text-type image in the image thumbnail in a highlighting manner relative to a scaled image.
[0090] The highlighting method can be flexibly set according to actual needs. For example, different font colors, font styles, and font distributions can be used to highlight text keywords relative to the scaled image. Specifically, when the terminal displays at least one text keyword associated with a text-type image in the image thumbnail, the text keyword is displayed according to the highlighting method relative to the scaled image, thereby highlighting the text keyword and making it easier for the user to identify.
[0091] In a specific application, such as Figure 4 As shown, in the image thumbnail Figure 2In the image, keywords are highlighted by bolding, underlining, and italics, and are scaled relative to the image size, so that users can accurately identify the keywords.
[0092] In this embodiment, the terminal highlights the text keywords associated with the text-type image according to the highlighting method relative to the scaled image, which helps users identify the text keywords associated with the text-type image.
[0093] In one embodiment, displaying at least one text keyword associated with a text-type image in an image thumbnail, in a highlighting manner relative to a scaled image, includes: displaying at least one text keyword associated with a text-type image in at least one text display area in the image thumbnail; and highlighting the font color of the at least one text keyword relative to the background color of the text display area where the at least one text keyword is located.
[0094] The text display area is the region within the image thumbnail where the user can display text keywords; specifically, it can be a text box. The text display area can display at least one text keyword; that is, it can display all or part of the text keywords associated with the text-type image. Font color refers to the color of the font when the text keywords are displayed. Background color refers to the color of the background of the text display area.
[0095] Specifically, the terminal divides at least one text display area within the image thumbnail, and displays at least one text keyword associated with the text-type image in each text display area. For example, a text display area can be divided within the image thumbnail, and all text keywords associated with the text-type image can be displayed in this text display area. Alternatively, the text display area in the image thumbnail can be the same as the text keyword area, with each text display area used to display one text keyword. The terminal displays at least one text keyword associated with the text-type image in the text display area. For the displayed text keyword, its font color is highlighted relative to the background color of the text display area. For example, the font color of the text keyword can be a complementary color to the background color of the text display area, thus highlighting it relative to the scaled image. For instance, when the background color of the text display area is white, the font color of the text keyword can be black; and when the background color of the text display area is black, the font color of the text keyword can be white, thus ensuring that the text keyword is highlighted relative to the scaled image.
[0096] In this embodiment, the font color of the text keywords displayed on the terminal is highlighted relative to the background color of the text display area where the text keywords are located, so as to ensure the recognizability of the text keywords and help users identify the text keywords associated with the text type image.
[0097] In one embodiment, the media preview method further includes: displaying a keyword operation area for the text-type image in response to a keyword editing trigger operation on the text-type image; and displaying at least one set keyword associated with the text-type image by the editing operation in response to an editing operation triggered in the keyword operation area.
[0098] The keyword editing trigger is an operation initiated by the user to edit the text keywords associated with the text-type image. The keyword operation area is the region where editing operations are performed on the text keywords associated with the text-type image. Editing operations include changing, adding, or deleting the text keywords associated with the text-type image. Setting keywords refers to the text keywords set by the user for the text-type image through editing operations.
[0099] Specifically, users can actively edit the text keywords associated with text-type images. For example, a user can trigger a keyword editing operation based on the text keywords in a text-type image. For instance, a user can access the attribute information of a text-type image and trigger a keyword editing operation based on the text keywords within that attribute information. Alternatively, to trigger a keyword editing operation based on text keywords in a text-type image, a user can long-press on a text keyword in the image to initiate the operation. In response to the user's keyword editing trigger, the terminal displays a keyword operation area for the text-type image, where the user can edit the text keywords.
[0100] The terminal responds to an editing operation triggered by the user in the keyword operation area, targeting at least one text keyword. It then displays the at least one keyword associated with the text-type image, set through the editing operation, thereby showing the result set by the user and allowing the user to confirm the editing operation. In specific implementations, users can modify, add, delete, or sort the text keywords of the text-type image to obtain at least one keyword associated with the text-type image.
[0101] Furthermore, in the image thumbnail, at least one text keyword associated with the text-type image is displayed, including: in the image thumbnail, at least one set keyword associated with the text-type image is displayed.
[0102] Specifically, the terminal displays the user-defined keywords associated with the text type images in the image thumbnails, so that the user can accurately distinguish between different text type images by setting the keywords.
[0103] In this embodiment, users can trigger editing operations on the text keywords associated with text-type images to personalize the text keywords associated with text-type images according to actual needs. The terminal displays the user-defined keywords associated with the text-type images, enabling users to accurately identify them based on the set keywords during preview. Furthermore, setting keywords increases the amount of information presented during media preview, which is beneficial for users to select the media resources they want to view based on the preview thumbnails.
[0104] In one embodiment, the media preview method further includes: displaying a full media display area in response to a triggering operation on a target keyword in at least one text keyword; and focusing on displaying a target text region in a text-type image within the full media display area; wherein the target text region is a text region in the text-type image associated with the content theme described by the target keyword.
[0105] The target keyword is the text keyword selected by the user from the displayed text keywords to trigger the display. The complete media display area is used to display the full media, where the original-size media can be shown in its entirety. The target text area is the text region within the text-type image associated with the content theme described by the target keyword; that is, the target keyword is identified based on the text within the target text area. Focusing the display on the target text area means concentrating the media display's focus on the target text area, i.e., displaying it with the target text area as the display focus. For example, the target text area can be displayed in the center.
[0106] Specifically, users can trigger actions based on displayed text keywords, such as clicking on a target keyword within the text keywords. In response to this action, the terminal displays the complete media display area, showcasing the text-type image within this area. Within this area, the terminal focuses on displaying the target text region within the text-type image, making it the focal point. For example, the terminal can center the display on the target text region. The target text region is the text area associated with the content theme described by the target keyword, allowing the terminal to quickly locate and focus on the text area associated with the target keyword.
[0107] In this embodiment, in response to the user's triggering operation on the target keyword in the text keywords, the terminal locates and focuses on the target text area in the text type image within the complete media display area. This focuses on the text area associated with the content theme described by the target keyword, allowing for quick location of the display focus according to the user's needs, simplifying media browsing operations and improving media browsing efficiency.
[0108] In one embodiment, displaying a media preview area of the media library in response to a preview trigger event for the media library includes: displaying an access point for the media library; and displaying the media preview area of the media library in response to a trigger operation for the access point.
[0109] The access point serves as the entry point for accessing the media library. Users can trigger actions through this access point to access and browse the various media within it. The access point can be configured according to actual needs, such as within the photo album interface or the local media cache interface. Specifically, the terminal displays the media library access point, and users can trigger actions through it. For example, if a user clicks the access point, the terminal responds by displaying a media preview area of the media library, where the user can preview the various media within the library.
[0110] Furthermore, the media preview method also includes: in response to a triggering operation on an image thumbnail, displaying the full image content of the text-type image in the media full display area of the text-type image.
[0111] The media full display area is used to display complete media, where the original size of the media can be displayed in its entirety. Complete image content refers to displaying text-type images in their original size; that is, displaying the original image of a text-type image.
[0112] Specifically, users can trigger operations on image thumbnails. For example, if a user clicks on an image thumbnail, the terminal responds to the user's trigger operation by displaying the full media display area of the text-type image. In this full media display area, the complete image content of the text-type image is displayed, that is, the terminal displays the text-type image at its original size.
[0113] In this embodiment, users can trigger a preview event for the media library through the media library access portal. By triggering the image thumbnail, the terminal displays the complete image content of the text-type image in the media full display area, thereby enabling users to access the complete media content in the media library.
[0114] In one embodiment, the media preview method further includes: displaying at least one text keyword associated with a text-type image in the full media display area; and, in response to a triggering operation on a target keyword among the at least one text keyword, focusing on displaying a target text region in the text-type image; wherein the target text region is a text region in the text-type image associated with the content theme described by the target keyword.
[0115] The target keyword is the text keyword selected by the user from the text keywords associated with the text-type image to trigger a specific action. The target text region is the text area within the text-type image associated with the content topic described by the target keyword; that is, the target keyword is identified based on the text within the target text region. Focusing the display on the target text region means concentrating the media display's focus on the target text region, i.e., displaying it with the target text region as the display focus. For example, the target text region can be displayed in the center.
[0116] Specifically, after displaying the complete image content of a text-type image within its media display area, the user can trigger an action based on associated text keywords. For example, the user can trigger the display of associated text keywords and select a target keyword. Responding to the user's action on the target keyword, the terminal focuses on displaying the target text area within the text-type image, making it the focal point. For instance, the terminal can center the display on the target text area. The target text area is the text area associated with the content theme described by the target keyword, allowing the terminal to quickly locate and focus on displaying the text area associated with the target keyword.
[0117] In this embodiment, when the terminal displays the complete image content of the text-type image in the media complete display area, it can also respond to the user's trigger operation on the target keyword in the text keywords associated with the text-type image, locate the target text area in the text-type image for focused display, so as to focus the text area associated with the content theme described by the target keyword. The focus of the display can be quickly located according to the user's needs, simplifying the operation of media browsing and improving the efficiency of media browsing.
[0118] In one embodiment, the partial area includes at least one of a text title area, a text body core area, or a text body highlight area in a text-type image. Displaying an image thumbnail pointing to a text-type image in the media preview area includes: displaying image thumbnails pointing to text-type images in the media preview area according to the order of the text-type images in the media library.
[0119] The text title area refers to the region in a text-type image that includes the title of the text; the text body core area refers to the region in a text-type image that includes the core content of the text body; and the text body highlighted area refers to the region in a text-type image that includes the highlighted content of the text body. A partial area may include at least one of the text title area, text body core area, or text body highlighted area in the text-type image. In practical applications, the text title area, text body core area, or text body highlighted area can be assigned priorities, such as the text title area having the highest priority, the text body core area second, and the text body highlighted area having the lowest priority. That is, the partial area selected for display in the image thumbnail after scaling is determined according to priority.
[0120] Specifically, the extracted region from the text-type image includes at least one of the text title region, the core text region, or the highlighted text region. In specific applications, regions can be determined according to priority from the text title region, the core text region, or the highlighted text region, and extracted from the text-type image to obtain a scaled image displayed in the image thumbnail based on these regions. The terminal determines the order of the text-type images in the media library, and in the media preview area, the terminal displays image thumbnails pointing to the text-type images according to their order in the media library, so that the order of each media in the media library is the same as the order of the image thumbnails pointing to the media in the media preview area. For example, the terminal can sort the text-type images according to their chronological order in the media library and display the image thumbnails pointing to those text-type images in an orderly manner in the media preview area.
[0121] In this embodiment, the "partial area" includes at least one of a text title area, a core text area, or a highlighted text area. This ensures that the text within the partial area is more important than the text outside the partial area. This allows the most important text areas in the text-type image to be displayed in the image thumbnail, enabling the high-importance text content to be directly displayed through the zoomed-in preview image. This increases the amount of information presented during media preview and helps users select the desired media resource based on the preview thumbnail. Furthermore, the terminal displays the corresponding image thumbnails according to the order of the text-type images in the media library, ensuring that the order of the image thumbnails in the media preview area matches the order in the media library, further facilitating user selection of the desired media resource based on the preview thumbnails.
[0122] In one embodiment, such as Figure 5 As shown, the media preview method also includes processing for determining text-type images, specifically including:
[0123] Step 502: Determine the target image in the media library.
[0124] The target image is an image in the media library that needs to be determined to be a text image. Specifically, the terminal determines the target image from the media library and then determines whether the target image is a text image. If the target image is a text image, a scaled-down image of a portion of the target image, scaled to a preset size, can be displayed in the image thumbnail.
[0125] Step 504: Perform character recognition on the target image to determine the text display area in the target image that includes text.
[0126] The text display area refers to the region in the target image that includes text. This text display area can be determined by character recognition of the target image. Specifically, the terminal can perform character recognition on the target image, such as using template matching character recognition algorithms, neural network character recognition algorithms, or support vector machine character recognition algorithms, to identify the text display area within the target image.
[0127] Step 506: When the proportion of the text display area in the target image reaches the text topic proportion threshold, the target image is determined to be a text type image.
[0128] Among them, the text theme proportion threshold is used to determine whether the target image belongs to the text type image. When the proportion of the text display area in the target image exceeds the text theme proportion threshold, it can be considered that the relationship between the text in the target image and the target image meets the text theme condition, that is, the target image has text as its theme content, and the terminal can determine that the target image belongs to the text type image.
[0129] Specifically, the terminal can determine the proportion of the text display area within the target image. In practice, the terminal can calculate this proportion based on the ratio of the text display area to the target image's total area, thus obtaining the text display area proportion. Alternatively, the terminal can also calculate the text display area proportion based on the ratio of the text display area to the non-background area of the target image. The terminal compares this text display area proportion with a text theme proportion threshold. If the text display area proportion exceeds the threshold, it indicates that the target image contains a large amount of text, and the text content is the main component of the image; therefore, the terminal determines that the target image is a text-type image. After determining that the target image is a text-type image, when previewing the target image, it is used as a text-type image in the image preview method. Specifically, a scaled-down image, where a portion of the target image is cropped and scaled to a preset size, is displayed in the image thumbnail.
[0130] In this embodiment, the text display area of the target image in the media library is determined by character recognition. When the proportion of the text display area reaches the threshold of the proportion of the text theme, the terminal determines that the target image belongs to the text type image. When the target image is previewed, it is treated as a text type image in the image preview method. This allows the important text content to be directly displayed through the zoomed-out preview image, increasing the amount of information presented during media preview and helping users select the media resources they want to view based on the preview thumbnails.
[0131] In one embodiment, when the proportion of the text display area in the target image reaches the text theme proportion threshold, it is determined that the target image belongs to the text type image, including: obtaining image source information of the target image; determining the text theme proportion threshold based on the image source information; and determining that the target image belongs to the text type image when the proportion of the text display area in the target image reaches the text theme proportion threshold.
[0132] Image source information refers to information related to the source of the target image, which may include, but is not limited to, the image acquisition method and the image source platform. Different text topic proportion thresholds can be set for different image sources to accurately determine whether an image is a text-based image. For example, for screenshots from chat interfaces or historical conversations of instant messaging applications, the main content is generally the text records in the chat, so the text topic proportion threshold can be set lower; while for images taken by a camera, the main content is generally the captured image, so the text topic proportion threshold can be set higher, thus ensuring accurate determination of the image type based on the image's source context. In practical implementation, the text topic proportion thresholds for images from different scenarios can be preset according to actual needs.
[0133] Specifically, the terminal acquires the image source information of the target image. This information can include the acquisition method, source platform, and other details about the scene from which the image originates. The image source information can be obtained from the target image's attribute information. Specifically, the terminal can query the target image's attribute information, which records various information related to the target image, such as image size, format, shooting device, and update time. The terminal retrieves the image source information of the target image from its attribute information. Based on the target image's image source information, the terminal determines a text theme proportion threshold. The mapping relationship between image source information and the text theme proportion threshold can be preset according to actual needs. The terminal can retrieve the text theme proportion threshold associated with the target image's image source information based on the preset mapping relationship between image source information and the text theme proportion threshold. The terminal compares the proportion of the text display area in the target image with the text theme proportion threshold. If the proportion of the text display area in the target image reaches the text theme proportion threshold, the target image is considered to contain a large amount of text content, and the target image is determined to be a text-type image.
[0134] In this embodiment, the terminal determines the text theme proportion threshold based on the image source information of the target image, and judges the proportion of the text display area in the target image based on the determined text theme proportion threshold to determine whether the target image belongs to the text type image. Thus, the type of the target image can be determined by combining the source scene of the target image, which can improve the accuracy of the target image type determination.
[0135] In one embodiment, the media preview method further includes: determining a text body region including the text from a text region including text in a text-type image; cropping a portion of the text body region; and obtaining a scaled image based on the portion of the text body region.
[0136] In this context, the text area refers to the region containing text within a text-type image, while the text body area is the region within the text-type image that includes the main text. The text in a text-type image includes text from various scenarios, such as article text, terminal status bar text, terminal notification bar text, or control text. The main text refers to the text related to the main content of the text-type image. For example, if the text-type image is a screenshot of a historical conversation, the main text could be the conversation message in the chat interface; if it's a screenshot of an article, the main text could be the text portion of the article; and if it's a screenshot of an announcement, the main text could be the text portion of the announcement. The main text in a text-type image is the text content that the user needs to obtain and can be used to distinguish between different text-type images. A portion of the image is extracted from the main text area, which can specifically include at least one of the text title area, the core text area, or the highlighted text area.
[0137] Specifically, the terminal can perform character recognition on text-type images to determine text regions containing text within the image, and then determine text-body regions containing the main text from these text regions. The terminal then crops a portion of the text-body region and obtains a scaled image displayed in the image thumbnail based on this cropped portion. In practical implementation, the terminal can perform character recognition on the text-type image to determine each text region within it, and further evaluate the text within each region, such as determining whether the text belongs to the main text based on its content characteristics, thereby determining the text-body regions containing the main text from each text region. The terminal can also crop a portion of the text with higher text importance based on the importance of each piece of text within the main text region, and obtain a scaled image based on this portion; for example, the portion can be scaled to obtain a scaled image.
[0138] In this embodiment, the terminal extracts a portion of the text with high text importance from the text-type image to obtain a scaled image. This allows the display of a portion of the text content in the image thumbnail through scaling, enabling the direct display of text content with high text importance through the previewed scaled image. This increases the amount of information presented during media preview and helps users select the media resources they wish to view based on the preview thumbnail.
[0139] In one embodiment, extracting a portion of the text from the text body region includes: extracting a text title region from the text body region whose text font size meets the title determination criteria, and obtaining a portion of the region based on the text title region.
[0140] The text font size refers to the font size of the text within the main text area. The title determination condition is used to determine whether text within the main text area is title text. Specifically, this can include a font size threshold; text exceeding the threshold is considered title text. The font size threshold can be preset with a specific value or adaptively set based on the font size of each text element within the main text area, thus accurately identifying title text within the main text area. The text title area refers to the region within the main text area that includes the title text.
[0141] Specifically, the terminal can determine the font size of each text element in the main text area and compare the font size of each text element with the title determination criteria. Based on the font size of each text element, the terminal can determine a text title area, including the text title, from the main text area. The terminal then obtains a partial area based on the text title area. For example, the terminal can extract the text title area from the main text area and use the extracted text title area as the text area to be previewed and displayed in the media preview area.
[0142] In this embodiment, based on the text title area in the text body area where the text font size meets the title determination criteria, the area that needs to be previewed in the media preview area is determined. This allows the text title to be displayed in the image thumbnail, presenting the text content with high importance through the title, increasing the amount of information presented during media preview, and helping users select the media resources they want to view based on the preview thumbnail.
[0143] In one embodiment, extracting a portion of the text from the main text region includes: extracting the core text region from the main text region where the text importance quantification parameter satisfies the core text determination condition, and obtaining a portion of the text based on the core text region.
[0144] The text importance quantification parameter is a quantitative parameter that characterizes the importance of text. Specifically, it can be a text importance score obtained by estimating the text's importance. Based on the text importance quantification parameter, the importance of each text in the main text area can be quantified, thereby determining the core text content within the main text area. The core text area can be a text region containing the core text content. The core text determination criterion is used to determine whether the text in the main text area belongs to the core text content. Specifically, it can include a quantification parameter threshold and the quantification parameter ranking result. For example, texts whose text importance quantification parameter reaches the quantification parameter threshold can be determined as core text content; alternatively, texts can be ranked according to their respective text importance quantification parameters, and the text with the highest text importance quantification parameter value can be determined as core text content.
[0145] Specifically, the terminal can determine the text importance quantification parameters of each text element in the main text area. These parameters can be obtained by analyzing each text element, such as by statistically analyzing the word frequencies of each text element and then using the statistical results to quantify the text importance parameters. In practice, the terminal can calculate the text importance quantification parameters of each text element in the main text area based on the TF-IDF (Term Frequency – Inverse Document Frequency) algorithm, specifically the TF-IDF score. The terminal then compares the text importance quantification parameters of each text element with the core text determination criteria to identify the core text area containing the core content. Based on these core text areas, the terminal obtains a partial area. For example, the terminal can extract the core text area from the main text area and use this extracted core text area as the text area to be previewed in the media preview area.
[0146] In this embodiment, based on the text importance quantification parameters in the text body area that meet the text core determination conditions, the part of the text that needs to be previewed in the media preview area is determined. This allows the core text content to be displayed in the image thumbnail, presenting text content with high text importance through the core text content. This increases the amount of information presented during media preview, which is beneficial for users to select the media resources they want to view based on the preview thumbnail.
[0147] In one embodiment, extracting a portion of a text area includes: extracting a highlighted text area from the text area, and obtaining a portion of the text area based on the highlighted text area.
[0148] Text highlighting refers to displaying text in a highlighted manner, such as using different font types, font sizes, font colors, underlines, italics, or bolding. The text body highlighting area is the text area that includes the highlighted text.
[0149] Specifically, the terminal can determine the display method of each text text in the text body area, thereby determining the highlighted text in each text body, and determining the text body highlighting area based on the highlighted text. The terminal obtains a partial area based on the text body highlighting area. For example, the terminal can extract the text body highlighting area from the text body area and use the extracted text body highlighting area as the text area that needs to be previewed and displayed in the media preview area.
[0150] In this embodiment, based on the highlighted text area in the main text area, the area that needs to be previewed in the media preview area is determined. This allows the highlighted text content to be displayed in the image thumbnail, presenting text content with high importance and increasing the amount of information presented during media preview. This helps users select the media resources they want to view based on the preview thumbnail.
[0151] In one embodiment, extracting a portion of a text area includes: determining a scaling method for the text area; scaling the text area according to the scaling method and a preset size to obtain a scaled text area; and cropping the scaled text area according to a preset size to obtain a portion of the text area whose size matches the preset size.
[0152] The scaling methods include scaling based on image height and scaling based on image width. When the height of the terminal display area is greater than its width, scaling based on image width can be used. The preset size can be flexibly set according to actual needs, and can be set to a square of M pixels.
[0153] Specifically, the terminal determines the scaling method for the text body area. This can be based on the relationship between the height and width of the terminal's display area, deciding whether to scale according to the image height or the image width. The terminal scales the text body area according to the determined scaling method and the preset size of the image thumbnail, so that the height or width of the text body area is scaled to the same as the preset size, resulting in a scaled text body area. In a practical implementation, if scaling according to the image width, the text body area can be scaled to the same width as the image thumbnail. The terminal then crops the scaled text body area according to the preset size of the image thumbnail, thereby extracting a portion of the area. For example, after scaling according to the image width to the same width as the image thumbnail, the terminal can crop the scaled text body area according to the height of the image thumbnail, extracting a portion with the same height as the image thumbnail. The width of the obtained partial area is the same as the width of the image thumbnail, and the height of the partial area is also the same as the height of the image thumbnail, that is, the size of the partial area is the same as the preset size of the image thumbnail.
[0154] In this embodiment, the terminal scales and crops the text area sequentially according to a determined scaling method and preset size to obtain a partial area whose size matches the preset size. The text within the partial area is more important than the text outside the partial area, so that the text content with high importance can be directly displayed through the zoomed preview image, which increases the amount of information presented during media preview and helps users select the media resources they want to view based on the preview thumbnails.
[0155] In one embodiment, the media preview method further includes: performing word segmentation on the text in the text-type image to obtain individual text segments; estimating the text importance of each text segment to obtain a text importance quantification parameter for each text segment; and determining at least one text keyword from each text segment based on the text importance quantification parameter.
[0156] Text segmentation involves dividing the text in a text-type image into individual words. The number of characters in each word is not fixed and can be one, two, or more. The text importance quantification parameter is a parameter that characterizes the importance of the text. Specifically, it can be a text importance score obtained by estimating the text's importance. Based on the text importance quantification parameter, the importance of each text element in the main text region can be quantified, thereby identifying the text keywords in the main text region. These text keywords are used to describe the content theme of the text in the text-type image.
[0157] Specifically, the terminal can perform word segmentation on the text in a text-type image to obtain individual text words. In practice, the terminal can use named entity recognition to segment the text in the text-type image, obtaining individual text words. The terminal can estimate the text importance of each text word, specifically using frequency statistics methods or pre-trained artificial neural network models or deep learning models, to obtain a quantitative parameter for the text importance of each text word. Based on the quantitative parameter, the terminal determines at least one text keyword from each text word. For example, the terminal can identify the at least one text word with the highest quantitative parameter value as the text keyword associated with the text-type image.
[0158] In this embodiment, after the terminal performs word segmentation on the text in the text-type image, it estimates the text importance of each text segment and determines at least one text keyword from each text segment based on the text importance quantification parameter of each text segment. Thus, determining text keywords from the text in the text-type image based on the text importance quantification parameter can ensure the accuracy of the text keywords.
[0159] In one embodiment, the media preview method further includes: establishing an association between at least one text keyword and a text-type image.
[0160] The association can be a mapping between text keywords and text-type images. For example, a mapping table can be used to record the association between text keywords and text-type images. Specifically, after the terminal determines the text keywords of a text-type image, it can establish at least one association between the text keywords and the text-type image and store the association so that the associated text keywords can be quickly retrieved based on the text-type image.
[0161] Furthermore, in the image thumbnail, displaying at least one text keyword associated with the text-type image includes: querying for at least one text keyword associated with the text-type image based on the association relationship, and displaying at least one text keyword in the image thumbnail.
[0162] Specifically, when a terminal displays text keywords for a text-type image, it can query the association relationships between the text-type images, determine at least one text keyword associated with the text-type image based on these relationships, and display the at least one queried text keyword in the image thumbnail. These association relationships can be pre-stored, so that when text keywords for a text-type image need to be displayed, the terminal can quickly query the associated text keywords based on these relationships, ensuring efficient display of the text keywords.
[0163] In this embodiment, the terminal establishes a relationship between text keywords and text-type images, so that it can quickly query the text keywords associated with the text-type images based on the relationship, which helps to improve the processing efficiency of text keyword display.
[0164] In one embodiment, the media preview method further includes: determining the text size of the text in the scaled image; obtaining the resolution of the text in the scaled image based on the font size; and determining the size relationship between the resolution of the text in the scaled image and a text visual resolution threshold.
[0165] The text resolution can be determined based on the font size, specifically the font height. For example, the font height can be directly used as the text resolution, thus the text visual resolution threshold includes the font height threshold. When the font height is greater than or equal to the font height threshold, the text resolution is considered greater than or equal to the text visual resolution threshold, meaning the text can be accurately recognized. The text visual resolution threshold can be preset according to actual needs and can also be customized by the user. For example, users can set the text visual resolution threshold based on their own recognition ability.
[0166] Specifically, after obtaining the scaled image, the terminal determines the font size of the text in the scaled image and, based on the font size, determines the resolution of the text in the scaled image. For example, the font height, which is included in the font size, can be determined as the text resolution. The terminal queries a pre-set text visual resolution threshold and compares the resolution of the text in the scaled image with the text visual resolution threshold to determine the size relationship between the resolution of the text in the scaled image and the text visual resolution threshold. If the resolution of the text in the scaled image is not less than the text visual resolution threshold, it indicates that the font size of the text in the scaled image is sufficient for the user to recognize, and the text in the scaled image can be determined to meet the text recognition conditions. Furthermore, for scaled images whose included text meets the text recognition conditions, the terminal may not display the text keywords associated with the text type image; while for scaled images whose included text does not meet the text recognition conditions, i.e., scaled images whose included text resolution is less than the text visual resolution threshold, the terminal may display the text keywords associated with the text type image, thereby ensuring that text content with high text importance in the text type image can be directly and effectively displayed.
[0167] In this embodiment, the terminal determines the relationship between the resolution of the text in the scaled image and the text visual resolution threshold based on the resolution determined by the font size of the text in the scaled image and a preset text visual resolution threshold. This allows for an accurate assessment of the recognition level of the text in the scaled image based on the font size and the text visual resolution threshold.
[0168] This application also provides an application scenario in which the media preview method described above is applied. Specifically, the media preview method is applied in this scenario as follows:
[0169] Currently, in applications that support image or video browsing, image lists are displayed using cropped thumbnails with reduced image resolution, allowing users to preview the overall content of the image. However, for images containing a large amount of text, such as announcements or text screenshots, the resolution compression results in the text in the thumbnails being too small to read. Furthermore, due to image thumbnails... Figure 1 Traditionally, images are displayed using square cropping, which easily loses information about the beginning and end of rectangular images. Users often struggle to preview the core content of an image through a thumbnail, requiring them to expand the image to understand the information, resulting in a poor user experience. Therefore, the media preview method provided in this embodiment extracts the main text area and identifies the location of key text information for text-based images. Then, it crops and scales the image to generate a thumbnail, prioritizing the display of the core text content and reducing the display of irrelevant information. If the thumbnail text is still smaller than the visible resolution, it further displays keywords related to the text in the image, ensuring that the thumbnail text can be recognized. Users can understand the key information of the text-based image through the thumbnail, improving the information expression capability of text-based images.
[0170] Thumbnails are small, scaled-down, and cropped images used for previewing pictures. Due to their small size, they load very quickly and are typically used for quick previews in lists. Resolution is a parameter that measures the amount of data within a bitmap image, usually expressed as pixels per inch (ppi) or dots per inch (dpi). The visual resolution of text refers to the minimum resolution at which the text content can be accurately recognized. A pixel is the smallest unit in an image represented by a sequence of numbers.
[0171] The media preview method provided in this embodiment can be applied to the thumbnail display of text-type images. By extracting the text content and position of the image, removing non-text areas, scaling the image, and then cropping it starting from the position of the key text, a thumbnail of the key area is obtained. This effectively improves the display efficiency of key information in the thumbnail and reduces the display of invalid information. When the thumbnail is still smaller than the visible resolution, the main text keywords are extracted and the text thumbnail content is displayed as floating keywords, effectively avoiding the problem of the thumbnail text being too large to read clearly, and making it easier for users to quickly understand the theme of the image, thus improving the user experience.
[0172] Specifically, such as Figure 6As shown, the media preview method provided in this embodiment can be applied to the scenario of browsing a terminal's photo album. In the terminal's photo album preview interface, thumbnails are displayed under the tiled image list. The images may contain text-based images, especially screenshots containing text that are frequently shared in social applications. Currently, the thumbnails reduce the image resolution, making it impossible to clearly see the text content. Only the approximate distribution of the text and the foreground and background colors can be seen. Users cannot intuitively preview the image text information through the thumbnails. Figure 7 As shown, the media preview method provided in this embodiment can also be applied to scenarios where applications browse acquired media. In the application, users can browse acquired images and videos, and also search for each media item. In the browsing interface, the terminal displays thumbnails of each media item. For images that are primarily text-based, the thumbnails are compressed, making the text in the thumbnails unrecognizable. Users need to click to view the original image to obtain detailed image information.
[0173] The current thumbnails are centered and compressed into squares by default, resulting in non-square areas of rectangular images not being displayed in the thumbnails. Furthermore, because images may contain non-text areas, invalid information may be displayed in the thumbnails. Therefore, the default thumbnail generation method is inefficient in conveying information. Figure 8 As shown, an image containing a holiday notice is a rectangular image. The media preview method provided in this embodiment displays its thumbnail as follows: Figure 9 As shown, the text area containing the holiday notice is enlarged and cropped into a square to be displayed as a thumbnail. When displayed, the text in the thumbnail meets the text recognition criteria, allowing users to directly obtain the important information from the original image. For example... Figure 10 As shown, for a rectangular article screenshot, since it is a rectangular image, directly centering, cropping, and compressing it into a thumbnail would not allow users to quickly obtain the key information of the article. The media preview method provided in this embodiment, when displaying its thumbnail, such as... Figure 11 As shown, the area containing the article's title is cropped and enlarged. When displayed, the text in the thumbnail meets the text recognition criteria, and users can directly obtain the article's title from the original image, thus quickly understanding the article's key information, specifically the theme of a nutritionist's guidance on a light diet. The media preview method provided in this embodiment identifies the main text area and confirms the location of key text areas, prioritizing the display of key text areas, allowing the thumbnail to more intuitively reflect the text's theme.
[0174] Furthermore, if the text resolution of the thumbnail after cropping key areas is still lower than the visible resolution, such as when there is a lot of text in the image that is difficult to read, keywords can be extracted from the image for display. For example... Figure 12 As shown, users can turn on the switch to display image keywords in the image manager, toggling the display of text keywords in images. Text-type image thumbnails display floating keyword text, which can be highlighted using a complementary color to the background color, allowing users to quickly preview key information. Users can long-press the thumbnail to trigger a full-screen display of the image content. For example, the first thumbnail displays keywords such as "index, US stocks, interest rate hike, large-cap, small-cap," thus describing the theme of the image pointed to by the thumbnail and facilitating a quick understanding of the key information.
[0175] Specifically, such as Figure 13As shown, the media preview method provided in this embodiment includes steps 1302 to 1320, wherein: Step 1302, triggering media preview; specifically, the user can click to access the media library and trigger the preview of each media in the media library; Step 1304, determining whether thumbnail cache exists; the terminal can determine whether the thumbnails corresponding to each media in the media library have been cached in advance. If so, it directly jumps to step 1320 to display each thumbnail for preview; Step 1306, if no thumbnail cache exists, extracting text from the image; the terminal can extract text from the image using character recognition technology. Text; Step 1308, determine if it is a text type image; the terminal can determine whether the image belongs to the text type image based on the text extracted from the image; if the image does not belong to the text type image, its corresponding thumbnail is directly displayed, which can be an image obtained by scaling the image; Step 1310, if the image belongs to the text type image, the terminal removes the non-text area content from the image; the terminal can remove content unrelated to the text content, such as status bar text, notification bar text, interface element text, etc., and retain the text area content; Step 1312, for the text area The terminal performs key text location recognition on the retained main text area to identify the location of key text within the main text area; Step 1314: Crop out the key area; After determining the location of key text in the main text area, the terminal crops out the key area containing the key text; Step 1316: Determine if the cropped key area is smaller than the visible resolution; If not, the terminal can scale the key area and display it as a thumbnail; Step 1318: If the cropped key area is smaller than the visible resolution, the terminal will display the text in the thumbnail... Keyword extraction is performed; in step 1320, a preview thumbnail is displayed; for non-text images, the displayed thumbnail can be an image obtained by directly scaling the original image; for images where the resolution of the key area is not less than the visible resolution, the displayed thumbnail can be an image obtained by scaling the key area; for images where the resolution of the key area is less than the visible resolution, the displayed thumbnail can be an image obtained by scaling the key area, and the keywords obtained through keyword extraction are displayed floating above the thumbnail, with the displayed keywords not less than the visible resolution.
[0176] Furthermore, the media preview method provided in this embodiment can be applied to images whose main content is text. The determination of a text-type image can be made by recognizing the proportion of the text area to the entire image. For example, text content can be extracted using OCR (Optical Character Recognition). OCR refers to the process of analyzing and recognizing image files of text materials to obtain text and layout information. The proportion of the text area in the entire non-background color area is calculated, with the background color being the most frequently occurring color in the image. When the proportion of the text area exceeds a certain threshold, such as 85%, the image can be labeled as a text-type image. Figure 14 As shown, for Figure 10 After extracting the text content from the article screenshot using OCR, it was determined that the text area in the article screenshot is the area covered by the rectangle. Based on the area ratio of the text area, it can be determined that the article screenshot belongs to the text-based image type.
[0177] Furthermore, image thumbnails can be generated by cropping and scaling proportionally with a centered position, prioritizing either width or height. For example... Figure 15 As shown, the original image's width is greater than its height. The image is cropped with height as the priority, centered. The final cropped thumbnail area is shown within the dashed box. Figure 16 As shown, the original image's height is greater than its width. It is cropped with width as the priority, centered, and the final cropped thumbnail area is shown in the dashed box. The media preview method provided in this embodiment extracts the text content and its location from the image using OCR. After removing non-text areas, it determines the location of key text information and uses the location of the key text area as a reference point for cropping. In the dynamic cropping of thumbnails of key text areas, as... Figure 17 As shown, the cropping is performed by centering the key area with a width priority. Specifically, the text area of the image is determined from the original image, and the starting position of the key text is determined within the text area. The image is then cropped according to the starting position of the key text at the size of the thumbnail. The final cropped thumbnail range is shown in the dashed box.
[0178] Since many text-based images are obtained through screenshots, these images often contain areas unrelated to the main text, such as the top status bar of a smartphone screenshot or blank areas at the beginning and end of notification images. Traditional thumbnail cropping includes these areas in the thumbnail preview, resulting in smaller and harder-to-read text. The media preview method provided in this embodiment uses OCR to identify all text content when cropping thumbnails. Based on the characteristics of non-textual text, such as the text in the top status bar of a screenshot (which typically contains timestamps and numbers) or button text in an image (which consists of only a few characters), only the main text area is cropped, increasing the effective text area in the thumbnail. Furthermore, to prevent the cropped text from being too close to the image edge, a very small portion of the non-textual text area can be retained for aesthetic purposes.
[0179] Furthermore, for the identification of key text in images obtained through OCR recognition, the methods for determining key text can include, but are not limited to, whether it is a text title or the importance of the text paragraph. Specifically, for the recognition of text titles in images, since titles are generally bolder and larger than regular fonts, the height of the text area obtained through OCR recognition can be used to determine the text size. Width cannot represent text size; the width of English characters and Chinese characters are not the same. A line with more than a certain threshold of consecutive characters and the largest text height can represent a first-level title. Thumbnails can be displayed primarily based on the title position. For paragraph importance, the text obtained through OCR recognition can be extracted, and then converted to full-width / half-width characters, simplified / traditional characters, etc., irrelevant to text importance can be removed. Punctuation marks, modal particles, adverbs, adjectives, numbers, etc., can be removed. The TF-IDF score is calculated for all words after word segmentation in each line. TF-IDF is a weighting technique used in information retrieval and data mining, where TF stands for Term Frequency and IDF stands for Inverse Document Frequency. TF-IDF calculation methods can include: calculating the term frequency (TF) of each word (number of times a word appears / total number of words in the image) and the inverse document frequency (IDF) (log(total number of documents in the corpus / number of documents containing the word + 1)), resulting in TF-IDF = TF * IDF. The importance score for each line of text is then obtained by summing and balancing the TF-IDF scores of all words, ultimately revealing the core paragraph of the text corresponding to the image. Furthermore, key text can be identified by recognizing emphasized text. Since emphasized text is typically highlighted by using colors other than body text, or by using italics or bold, the importance score of the text in those positions can be weighted by statistically analyzing the position and number of emphasized text.
[0180] The specific thumbnail cropping process includes: First, determining whether the thumbnail scaling method is height-first or width-first. Since mobile terminals generally have a height greater than their width, width-first scaling is usually preferred, so we'll take width-first scaling as an example. After removing non-text content from the image, scale the image. The thumbnail scaling ratio = text area width / thumbnail width. The scaled image is then cropped. The starting horizontal and vertical coordinates for cropping are [0 (width-first scaling has a horizontal coordinate of 0), starting height of key text / scaling ratio]. The vertical coordinate cannot exceed the height of the scaled image minus the thumbnail height. The cropped width and height are the fixed sizes of the current thumbnail.
[0181] Since generating thumbnails requires scaling the image, even after cropping the key areas, the text in the thumbnail may still be too large to read after compression. In this case, the height of the text recognized by the OCR in the original image can be averaged across all text. After scaling, this averaged height is divided by the scaling ratio to obtain the thumbnail's text height, which represents the text resolution. When the thumbnail text height is less than a certain threshold (below the resolution of normal human vision), a keyword preview can be generated for the corresponding text content. Keyword generation methods can include, but are not limited to, traditional algorithms based on text importance such as TF-IDF, TextRank, and LDA (Latent Dirichlet Allocation) to mine important keywords, as well as keyword generation algorithms based on deep learning models.
[0182] Specifically, the text identified by OCR in the image is first segmented into words. This segmentation is then used to introduce Named Entity Recognition (NER) to extract names of people, places, organizations, and proper nouns from the image text, resulting in a final list of all words. Named Entity Recognition, also known as proper noun recognition, identifies entities with specific meanings in text, primarily including names of people, places, organizations, and proper nouns. A TF-IDF-based method can be used to calculate the TF-IDF score for each word using the formula above. Words are then sorted in descending order of their TF-IDF scores, and the top N values can be used as thumbnail keywords. N can also be dynamically calculated based on whether the total length of the keywords exceeds the visible resolution.
[0183] Deep learning-based algorithms can be categorized as either supervised or unsupervised. Supervised methods extract text features such as TF-IDF values, first-occurrence positions, presence in titles, parts of speech, and contextual features. These features are then combined with training corpora and trained using different neural network architectures to obtain a prediction model. The prediction model yields the keyword probability for each word, and the importance probability of each keyword is obtained by reversing the model's prediction values. Unsupervised methods, such as using a large-scale pre-trained text model like BERT (Bidirectional Encoder Representation from Transformers), extract document embeddings and embeddings for all words. BERT is a pre-training technique used in Natural Language Processing (NLP). The similarity between candidate keywords and documents is calculated, and the top N words are selected as the final keywords by reversing the similarity ranking.
[0184] After generating thumbnail keywords, they can be displayed floating above the thumbnail in the media preview interface. Users can choose whether to enable this feature, and long-pressing the thumbnail will display the full image corresponding to the thumbnail. Furthermore, to avoid users repeatedly calculating and judging the same information when viewing images later, the processed results can be cached locally. The next time a user views an image, the cached result will be read first and rendered to generate a thumbnail, improving the speed at which users view preview thumbnails.
[0185] The media preview method provided in this embodiment increases the effective text information content of thumbnails by scaling and cropping key areas of text-type images, and optimizes the display of key information in the article, allowing users to easily understand the main content of the image through the thumbnail. When the text resolution is still lower than the user's visual resolution, keywords are displayed in a floating manner, avoiding the problem of thumbnails being unclear when there is a lot of text. This solution can effectively improve the information dissemination efficiency of thumbnails for text-type images, allowing users to quickly preview the core content of images in the image list through thumbnails.
[0186] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0187] Based on the same inventive concept, this application also provides a media preview device for implementing the media preview method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more media preview device embodiments provided below can be found in the limitations of the media preview method described above, and will not be repeated here.
[0188] In one embodiment, such as Figure 18 As shown, a media preview device 1800 is provided, including: a preview area display module 1802, a thumbnail display module 1804, and a text area display module 1806, wherein:
[0189] The preview area display module 1802 is used to display the media preview area of the media library in response to a preview trigger event for the media library;
[0190] The thumbnail display module 1804 is used to display an image thumbnail pointing to the text type image in the media preview area when the media library includes text type images; the image thumbnail has a preset size, and the proportion of the text display area in the text type image reaches the text theme proportion threshold.
[0191] The text region display module 1806 is used to display a scaled image of a portion of the text-type image, which is scaled to a preset size, in an image thumbnail pointing to a text-type image; the resolution of the text in the scaled image is not less than the text visual resolution threshold and meets the text recognition conditions; in the text-type image, the text within a portion of the image is more important than the text outside that portion of the image.
[0192] In one embodiment, a keyword display module is further included, which displays at least one text keyword associated with the text-type image in the image thumbnail; the at least one text keyword is used to describe the content theme of the text in the text-type image.
[0193] In one embodiment, the keyword display module is further configured to, in response to a keyword display trigger operation in the media preview area, display at least one text keyword associated with the text type image from which the zoomed image originates in the zoomed image.
[0194] In one embodiment, the system further includes an activation entry display module, configured to display a keyword display activation entry indicating an activation state in the media preview area; the keyword display module is further configured to, in response to a trigger operation on the keyword display activation entry, switch the display of the keyword display activation entry to indicate an activation state, and display at least one text keyword associated with the text type image from which the zoomed image originates in the zoomed image; the activation entry display module is further configured to, in response to a trigger operation on the keyword display activation entry indicating an activation state, switch the display of the keyword display activation entry to indicate an activation state, and hide at least one text keyword in the zoomed image.
[0195] In one embodiment, the keyword display module is further configured to display a preset number of text keywords associated with the text-type image in the image thumbnail; the preset number of text keywords are displayed in order of their respective text importance.
[0196] In one embodiment, the keyword display module is further configured to display at least one text keyword associated with the text type image in the image thumbnail when the resolution of the text in the scaled image is less than a text visual resolution threshold.
[0197] In one embodiment, the keyword display module is further configured to display at least one text keyword associated with the text-type image in the image thumbnail, in a highlighting manner relative to the scaled image.
[0198] In one embodiment, the keyword display module is further configured to display at least one text keyword associated with the text type image in at least one text display area in the image thumbnail; and to highlight the font color of the at least one text keyword relative to the background color of the text display area where the at least one text keyword is located.
[0199] In one embodiment, the system further includes a keyword editing trigger module, configured to display a keyword operation area for the text-type image in response to a keyword editing trigger operation on the text-type image; and to display at least one set keyword associated with the text-type image as set by the editing operation in response to an editing operation on at least one text keyword triggered in the keyword operation area; the keyword display module is further configured to display at least one set keyword associated with the text-type image in the image thumbnail.
[0200] In one embodiment, the system further includes a complete display area module and a focused display module; wherein: the complete display area module is used to display the complete display area of the media in response to a triggering operation of a target keyword in at least one text keyword; the focused display module is used to locate and focus on a target text area in a text-type image within the complete display area of the media; the target text area is a text area in a text-type image associated with the content theme described by the target keyword.
[0201] In one embodiment, the preview area display module 1802 is further configured to display an access point to the media library; in response to a trigger operation on the access point, display a media preview area of the media library; and further includes a thumbnail trigger response module, configured to display the complete image content of the text-type image in the media full display area of the text-type image in response to a trigger operation on the image thumbnail.
[0202] In one embodiment, the thumbnail trigger response module is further configured to display at least one text keyword associated with the text type image in the full media display area; in response to a trigger operation on a target keyword among the at least one text keyword, to locate and focus on display a target text area in the text type image; the target text area is a text area in the text type image associated with the content theme described by the target keyword.
[0203] In one embodiment, the partial area includes at least one of the text title area, the core text area, or the highlighted text area in the text type image; the thumbnail display module 1804 is further configured to display image thumbnails pointing to the text type images in the media preview area according to the order of the text type images in the media library.
[0204] In one embodiment, the system further includes a target image determination module, a character recognition module, and an image type determination module; wherein: the target image determination module is used to determine a target image in the media library; the character recognition module is used to perform character recognition on the target image and determine the text region in the target image that includes text; and the image type determination module is used to determine that the target image belongs to the text type image when the proportion of the text display area in the target image reaches the text theme proportion threshold.
[0205] In one embodiment, the image type determination module includes a source information determination module, a proportion threshold determination module, and a proportion comparison module; wherein: the source information determination module is used to obtain image source information of the target image; the proportion threshold determination module is used to determine the text topic proportion threshold based on the image source information; and the proportion comparison module is used to determine that the target image belongs to the text type image when the proportion of the text display area in the target image reaches the text topic proportion threshold.
[0206] In one embodiment, the image further includes a scaling image acquisition module, configured to determine a text body region including the text from a text region including text in a text-type image; extract a portion of the text body region; and obtain a scaling image based on the portion of the text body region.
[0207] In one embodiment, the scaling image acquisition module is further configured to: extract a text title region from the text body region where the text font size satisfies the title determination condition, and obtain a partial region based on the text title region; extract a text core region from the text body region where the text importance quantification parameter satisfies the core text determination condition, and obtain a partial region based on the text core region; and extract a text highlighting region from the text body region, and obtain a partial region based on the text highlighting region.
[0208] In one embodiment, the image scaling module is further configured to determine the scaling method for the text body region; scale the text body region according to the scaling method and the preset size to obtain the scaled text body region; and crop the scaled text body region according to the preset size to obtain a portion of the region whose size matches the preset size.
[0209] In one embodiment, the system further includes a segmentation processing module, an importance estimation module, and an estimation result processing module; wherein: the segmentation processing module is used to perform word segmentation processing on the text in the text-type image to obtain each text word; the importance estimation module is used to estimate the text importance of each text word to obtain the text importance quantification parameter of each text word; and the estimation result processing module is used to determine at least one text keyword from each text word based on the text importance quantification parameter.
[0210] In one embodiment, the system further includes an association establishment module for establishing an association between at least one text keyword and a text-type image; and a keyword display module for querying at least one text keyword associated with the text-type image based on the association, and displaying at least one text keyword in the image thumbnail.
[0211] In one embodiment, the system further includes a text recognition and determination module, used to determine the font size of the text in the scaled image; obtain the resolution of the text in the scaled image based on the font size; and determine the size relationship between the resolution of the text in the scaled image and a text visual resolution threshold.
[0212] Each module in the aforementioned media preview device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the operations corresponding to each module.
[0213] In one embodiment, a computer device is provided, which may be a terminal or a server, and its internal structure diagram may be as follows: Figure 19 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a media preview method. The display unit of the computer device is used to form a visually visible image. It can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0214] Those skilled in the art will understand that Figure 19 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0215] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0216] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0217] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0218] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0219] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0220] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification. The above embodiments only illustrate several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of this application. It should be noted that for those skilled in the art, several modifications and improvements can be made without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A media preview method, characterized in that, The method includes: In response to a preview trigger event for the media library, display the media preview area of the media library; When the media library includes text-type images, an image thumbnail pointing to the text-type image is displayed in the media preview area; the image thumbnail has a preset size, and the proportion of the text display area in the text-type image reaches a text theme proportion threshold; wherein, a portion of the text-type image is cropped and scaled to the preset size to obtain a scaled image; in the text-type image, the text within the portion of the image is more important than the text outside the portion of the image. If the resolution of the text in the scaled image is not less than the text visual resolution threshold, the scaled image is displayed in the image thumbnail; If the resolution of the text in the scaled image is less than the text visual resolution threshold, the scaled image is displayed in the image thumbnail, and at least one text keyword associated with the text type image is displayed in the scaled image; the at least one text keyword is used to describe the content theme of the text in the text type image; the resolution of the text keyword is not less than the text visual resolution threshold.
2. The method according to claim 1, characterized in that, The provision that at least one text keyword associated with the text type image is displayed in the image thumbnail includes: In response to a keyword display triggering operation in the media preview area, at least one text keyword associated with the text type image from which the zoomed image originates is displayed in the zoomed image.
3. The method according to claim 2, characterized in that, The method further includes: The activation entry point is displayed in the media preview area using keywords indicating the activation status. In response to a keyword display trigger operation in the media preview area, at least one text keyword associated with the text type image from which the zoomed image originates is displayed in the zoomed image, including: In response to the triggering operation of the keyword display activation entry, the keyword display activation entry is switched to display the activation status, and at least one text keyword associated with the text type image from which the zoomed image originates is displayed in the zoomed image; The method further includes: In response to a trigger operation that displays an activation entry for the keyword indicating an active state, the keyword activation entry is switched to display an active state, and the at least one text keyword is hidden in the zoomed image.
4. The method according to claim 1, characterized in that, The provision that at least one text keyword associated with the text type image is displayed in the image thumbnail includes: The image thumbnail displays a preset number of text keywords associated with the text type image; The preset number of text keywords are displayed in order of their respective textual importance.
5. The method according to claim 1, characterized in that, The provision that at least one text keyword associated with the text type image is displayed in the image thumbnail includes: In the image thumbnail, at least one text keyword associated with the text type image is displayed in a highlighted manner relative to the scaled image.
6. The method according to claim 5, characterized in that, The step of displaying at least one text keyword associated with the text type image in the image thumbnail, in a highlighting manner relative to the scaled image, includes: In at least one text display area of the image thumbnail, at least one text keyword associated with the text type image is displayed respectively; The font color of the at least one text keyword is highlighted relative to the background color of the text display area where the at least one text keyword is located.
7. The method according to claim 1, characterized in that, The method further includes: In response to a keyword editing trigger operation on the text-type image, a keyword operation area for the text-type image is displayed; In response to an editing operation on the at least one text keyword triggered in the keyword operation area, at least one set keyword associated with the text type image set by the editing operation is displayed; The provision that at least one text keyword associated with the text type image is displayed in the image thumbnail includes: The image thumbnail displays at least one set keyword associated with the text type image.
8. The method according to claim 1, characterized in that, The method further includes: In response to a triggering operation on the target keyword in the at least one text keyword, the entire media display area is displayed; Within the complete media display area, the target text region in the text-type image is located and displayed in focus. The target text region is the text region in the text type image associated with the content topic described by the target keyword.
9. The method according to claim 1, characterized in that, The step of displaying a media preview area of the media library in response to a preview trigger event for the media library includes: Displays the access point to the media library; In response to a triggered operation on the access point, the media preview area of the media library is displayed; The method further includes: In response to a triggering operation on the image thumbnail, the complete image content of the text-type image is displayed in the media full display area of the text-type image.
10. The method according to claim 9, characterized in that, The method further includes: In the media full display area, at least one text keyword associated with the text type image is displayed; In response to a triggering operation on a target keyword in the at least one text keyword, the target text region in the text type image is located and displayed in focus; The target text region is the text region in the text type image associated with the content topic described by the target keyword.
11. The method according to any one of claims 1 to 10, characterized in that, The partial area includes at least one of the text title area, the core text area, or the highlighted text area in the text type image; The step of displaying an image thumbnail pointing to the text type image in the media preview area includes: In the media preview area, image thumbnails pointing to the text type images are displayed according to their order in the media library.
12. The method according to claim 1, characterized in that, The method further includes: Identify the target image in the media library; Character recognition is performed on the target image to determine the text display area in the target image that includes text; When the proportion of the text display area in the target image reaches the text topic proportion threshold, the target image is determined to be a text type image.
13. The method according to claim 12, characterized in that, The step of determining that the target image is a text-type image when the proportion of the text display area in the target image reaches the text topic proportion threshold includes: Obtain the image source information of the target image; Determine the text topic proportion threshold based on the image source information; When the proportion of the text display area in the target image reaches the text topic proportion threshold, the target image is determined to be a text type image.
14. The method according to claim 1, characterized in that, The method further includes: From the text region containing text in the text type image, determine the text body region including the text body; Extract the portion of the text from the main text area; The scaled image is obtained based on the aforementioned partial region.
15. The method according to claim 14, characterized in that, The step of extracting the portion of the text from the main text area includes at least one of the following: From the text body area, extract the text title area whose font size meets the title determination criteria, and obtain the partial area based on the text title area; From the text body region, extract the text body core region where the text importance quantification parameter satisfies the text body core determination condition, and obtain the partial region based on the text body core region; From the text body area, extract the highlighted text body area, and obtain the partial area based on the highlighted text body area.
16. The method according to claim 14, characterized in that, The step of extracting the portion of the text from the main text area includes: Determine the scaling method for the text body area; The text area is scaled according to the scaling method and the preset size to obtain the scaled text area. The scaled text area is cropped according to the preset size to obtain a portion of the area whose size matches the preset size.
17. The method according to claim 1, characterized in that, The method further includes: The text in the text-type image is segmented to obtain individual text words; The text importance of each text segment is estimated to obtain the text importance quantification parameters of each text segment. Based on the text importance quantification parameters, at least one text keyword is determined from each text segmentation.
18. The method according to claim 17, characterized in that, The method further includes: Establish a relationship between the at least one text keyword and the text type image; The provision that at least one text keyword associated with the text type image is displayed in the image thumbnail includes: Based on the association, at least one text keyword associated with the text type image is retrieved, and the at least one text keyword is displayed in the image thumbnail.
19. The method according to any one of claims 1 to 10 or 12 to 18, characterized in that, The method further includes: Determine the font size of the text in the scaled image; The resolution of the text in the scaled image is obtained based on the font size; Determine the relationship between the resolution of the text in the scaled image and the text visual resolution threshold.
20. A media preview device, characterized in that, The device includes: The preview area display module is used to display the media preview area of the media library in response to a preview trigger event for the media library; A thumbnail display module is used to display an image thumbnail pointing to the text-type image in the media preview area when the media library includes text-type images; the image thumbnail has a preset size, and the proportion of the text display area in the text-type image reaches a text theme proportion threshold; wherein a portion of the text-type image is cropped and scaled to the preset size to obtain a scaled image; in the text-type image, the text within the portion of the image is more important than the text outside the portion of the image; A text area display module is used to display the scaled image in the image thumbnail when the resolution of the text in the scaled image is not less than the text visual resolution threshold; A keyword display module is used to display the scaled image in the image thumbnail when the resolution of the text in the scaled image is less than the text visual resolution threshold, and to display at least one text keyword associated with the text type image in the scaled image; the at least one text keyword is used to describe the content theme of the text in the text type image; the resolution of the text keyword is not less than the text visual resolution threshold.
21. The media preview device according to claim 20, characterized in that, The keyword display module is also configured to, in response to a keyword display trigger operation in the media preview area, display at least one text keyword associated with the text type image from which the zoomed image originates in the zoomed image.
22. The media preview device according to claim 21, characterized in that, The device also includes an activation entry display module, used to display keywords indicating the activation status in the media preview area to show the activation entry; The keyword display module is also used to switch the keyword display activation entry to an active state in response to the triggering operation of the keyword display activation entry, and to display at least one text keyword associated with the text type image from which the zoomed image originates in the zoomed image. The activation entry display module is also used to respond to a trigger operation of displaying the activation entry of the keyword indicating the activation status, switch the display of the keyword activation entry to indicate the pending activation status, and hide the at least one text keyword in the zoomed image.
23. The media preview device according to claim 20, characterized in that, The keyword display module is also used to display a preset number of text keywords associated with the text type image in the image thumbnail; the preset number of text keywords are displayed in order of their respective text importance.
24. The media preview device according to claim 20, characterized in that, The keyword display module is also used to display at least one text keyword associated with the text type image in the image thumbnail, in a highlighting manner relative to the scaled image.
25. The media preview device according to claim 24, characterized in that, The keyword display module is also used to display at least one text keyword associated with the text type image in at least one text display area in the image thumbnail; the font color of the at least one text keyword is highlighted relative to the background color of the text display area where the at least one text keyword is located.
26. The media preview device according to claim 20, characterized in that, The device further includes a keyword editing trigger module, which displays a keyword operation area for the text type image in response to a keyword editing trigger operation on the text type image; In response to an editing operation on the at least one text keyword triggered in the keyword operation area, at least one set keyword associated with the text type image set by the editing operation is displayed; The keyword display module is also used to display at least one set keyword associated with the text type image in the image thumbnail.
27. The media preview device according to claim 20, characterized in that, The device also includes a complete display area module and a focused display module; The complete display area module is used to display the complete media display area in response to a trigger operation on the target keyword in the at least one text keyword; The focused display module is used to locate and focus on a target text region in the text-type image within the complete media display area; the target text region is a text region in the text-type image associated with the content theme described by the target keyword.
28. The media preview device according to claim 20, characterized in that, The preview area display module is also used to display the access point of the media library; in response to a trigger operation on the access point, the media preview area of the media library is displayed; The device further includes a thumbnail trigger response module, which, in response to a trigger operation on the image thumbnail, displays the complete image content of the text-type image in the media full display area of the text-type image.
29. The media preview device according to claim 28, characterized in that, The thumbnail trigger response module is also used to display at least one text keyword associated with the text type image in the full display area of the media; in response to a trigger operation on a target keyword among the at least one text keyword, to locate and focus on display a target text area in the text type image; the target text area is a text area in the text type image associated with the content theme described by the target keyword.
30. The media preview device according to any one of claims 20 to 29, characterized in that, The partial area includes at least one of the text title area, the core text area, or the highlighted text area in the text type image; The thumbnail display module is also used to display image thumbnails pointing to the text type images in the media preview area, according to the sorting of the text type images in the media library.
31. The media preview device according to claim 20, characterized in that, The device also includes a target image determination module, a character recognition module, and an image type determination module; The target image determination module is used to determine the target image in the media library; The character recognition module is used to perform character recognition on the target image and determine the text display area in the target image that includes text; The image type determination module is used to determine that the target image is a text type image when the proportion of the text display area in the target image reaches the text topic proportion threshold.
32. The media preview device according to claim 31, characterized in that, The image type determination module includes a source information determination module, a proportion threshold determination module, and a proportion comparison module; The source information determination module is used to obtain the image source information of the target image; The proportion threshold determination module is used to determine the text topic proportion threshold based on the image source information; The proportion comparison module is used to determine that the target image is a text-type image when the proportion of the text display area in the target image reaches the text theme proportion threshold.
33. The media preview device according to claim 20, characterized in that, The device further includes a scaled image acquisition module, configured to determine a text body region including the text body from a text region including text in the text type image; to crop the partial region from the text body region; and to obtain the scaled image based on the partial region.
34. The media preview device according to claim 33, characterized in that, The zoomed image acquisition module is further configured to: extract a text title region from the text body region where the text font size satisfies the title determination condition, and obtain the partial region based on the text title region; extract a text core region from the text body region where the text importance quantification parameter satisfies the core text determination condition, and obtain the partial region based on the text core region. From the text body area, extract the highlighted text body area, and obtain the partial area based on the highlighted text body area.
35. The media preview device according to claim 33, characterized in that, The image scaling module is further configured to determine the scaling method for the text body region; and scale the text body region according to the scaling method and the preset size to obtain the scaled text body region. The scaled text area is cropped according to the preset size to obtain a portion of the area whose size matches the preset size.
36. The media preview device according to claim 20, characterized in that, The device also includes a word segmentation processing module, an importance estimation module, and an estimation result processing module; The word segmentation module is used to segment the text in the text type image to obtain individual text words; The importance estimation module is used to estimate the text importance of each text segment and obtain the text importance quantification parameters of each text segment. The estimation result processing module is used to determine at least one text keyword from each text segmentation based on the text importance quantification parameter.
37. The media preview device according to claim 36, characterized in that, The device further includes an association establishment module for establishing an association relationship between the at least one text keyword and the text type image; The keyword display module is also used to query at least one text keyword associated with the text type image based on the association relationship, and display the at least one text keyword in the image thumbnail.
38. The media preview device according to any one of claims 20 to 29 or 31 to 37, characterized in that, The device further includes a text recognition and determination module, used to determine the font size of the text in the scaled image; obtain the resolution of the text in the scaled image based on the font size; and determine the size relationship between the resolution of the text in the scaled image and the text visual resolution threshold.
39. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 19.
40. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 19.
41. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 19.
Citation Information
Patent Citations
Thumbnail generation method and system
CN103902730A
Method and system for displaying text in thumbnail
CN106951226A