Display content and data processing method, device, system and chip

By generating enhanced images based on user gaze duration and interaction metadata, and combining cloud servers and multimodal large models, the problem of high production cost and poor versatility of e-book AR/VR media is solved. This achieves low-threshold and efficient image enhancement effects, improving user experience and fun.

CN122018742APending Publication Date: 2026-05-12HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-08-06
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies, the production cost of AR/VR media for e-books is high and their versatility is poor, making it difficult to improve the reading experience.

Method used

By acquiring the user's gaze duration on the illustration, enhanced images are generated based on interactive metadata, supporting global or local interaction modes. Image enhancement data is provided by cloud servers, reducing the processing pressure on electronic devices. Furthermore, the main element features are identified through a multimodal large model, generating diverse enhanced images.

Benefits of technology

It achieves low-cost, low-barrier image enhancement, improves user interaction experience and reading enjoyment, and is applicable to various digital content formats, including digital picture books, digital catalogs and e-books, enhancing user immersion and realism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122018742A_ABST
    Figure CN122018742A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a display content and data processing method, device and system and a chip. The display content processing method comprises the steps that a first illustration and first interaction metadata are obtained, and the first interaction metadata is used for indicating generation of a first enhanced image corresponding to the first illustration; displaying the first illustration on a display screen; obtaining a first watching duration of the user on the first illustration; and when it is determined that the first fixation duration reaches the first preset duration, generating and displaying a first enhanced image according to the first interaction metadata. According to the display content processing method, the first enhanced image corresponding to the first illustration can be generated and displayed based on the first interaction metadata corresponding to the first illustration under the condition of determining that the user has a gazing intention on the first illustration, so that the display effect of the first illustration can be enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of digital reading technology, and in particular to a method, device, system, and chip for displaying content and processing data. Background Technology

[0002] Digital reading offers immense convenience to users due to its massive download volume, portability, and anytime accessibility. However, abstract and obscure text content and static illustrations cannot satisfy users' more diverse reading needs. Therefore, some manufacturers are developing AR / VR media for e-books using Augmented Reality (AR) or Virtual Reality (VR) technologies. When users read e-books, loading the corresponding AR / VR media and displaying the corresponding AR / VR animations enriches the digital reading experience, reduces comprehension difficulty, and especially enhances the fun of digital reading for children, helping to stimulate children's reading interest.

[0003] However, developing AR / VR media to accompany e-books requires specialized production tools and teams. The entire production process is lengthy and costly. Furthermore, terminal devices need specific AR / VR engines and hardware chips to load AR / VR media, making it difficult to guarantee universality in applications.

[0004] Therefore, how to ensure a good reading experience while reducing production costs and application difficulty has become a pressing technical problem to be solved in this field. Summary of the Invention

[0005] This application provides a method, device, system, and chip for displaying content and processing data to address the problems of high cost in producing image enhancement data and high barriers to entry in enhancing image display.

[0006] In a first aspect, embodiments of this application provide a display content processing method applied to an electronic device. The method includes: acquiring a first illustration and first interactive metadata, wherein the first interactive metadata is used to indicate the generation of a first enhancement; displaying the first illustration on a display screen; acquiring a first gaze duration of a user on the first illustration; and, if it is determined that the first gaze duration has reached a first preset duration, generating and displaying the first enhanced image based on the first interactive metadata.

[0007] In this embodiment of the application, based on the user's gaze point, when it is determined that the user has a gaze intention towards the first illustration, the first enhanced image corresponding to the first illustration can be directly generated and displayed according to the first interaction metadata. For electronic devices, there is no need to perform complex processing operations to enhance the image of the first illustration, which not only improves the display effect of the first illustration, but also has a low threshold for use and is more versatile.

[0008] In one optional embodiment, the first interaction metadata includes interaction object information, the interaction object information including the coordinate range corresponding to an independent interaction object in the first illustration, the interaction object information being used to indicate the interaction modes supported by the first illustration, wherein the interaction mode is a global interaction mode or a local interaction mode.

[0009] In this embodiment, based on the interactive object information included in the first interactive metadata, the independent interactive objects in the first illustration and the interactive modes supported by the first illustration can be determined. Based on this, different interactive operation methods can be provided to the user, which helps to improve the user experience.

[0010] In an optional embodiment, when it is determined that the first gaze duration has reached a first preset duration, the display content processing method further includes: if the coordinate range is the first coordinate range corresponding to the first illustration, determining that the first illustration supports a global interaction mode; in the case of the global interaction mode, displaying interactive prompt information corresponding to the first illustration on the display screen to prompt the user to perform interactive operations on the first illustration.

[0011] In an optional embodiment, when it is determined that the first gaze duration reaches a first preset duration, the display content processing method further includes: if the coordinate range is a second coordinate range corresponding to at least one sub-region or at least one main element in the first illustration, determining that the first illustration supports a local interaction mode; in the case of the local interaction mode, displaying interactive prompt information corresponding to at least one sub-region in the first illustration on the display screen to prompt the user to perform an interactive operation on the first sub-region, wherein the first sub-region is one of the at least one sub-region; or, in the case of the local interaction mode, displaying interactive prompt information corresponding to at least one main element in the first illustration on the display screen to prompt the user to perform an interactive operation on the first main element, wherein the first main element is one of the at least one main element.

[0012] In this embodiment, based on the interaction mode supported by the first illustration, interactive prompts corresponding to the corresponding interaction mode are displayed, which can guide users to perform interactive operations on the first illustration, its sub-regions, or main elements. This not only reduces the probability of users missing operations, but also provides a better user experience through personalized interactive prompts.

[0013] In an optional embodiment, generating and displaying the first enhanced image based on the first interaction metadata includes: obtaining image enhancement data corresponding to the first illustration based on the user's interaction operation on the target interactive object, wherein the target interactive object is one of the first illustration, a first sub-region in the first illustration, or a first main element; and generating and displaying the first enhanced image based on the first interaction metadata and the image enhancement data.

[0014] In this embodiment, the first illustration supports different interaction modes, allowing users to interact with the first illustration, its sub-regions, or main elements. This provides users with richer interaction options, enhances the interactive experience, and encourages user engagement.

[0015] In an optional embodiment, obtaining image enhancement data corresponding to the first illustration based on the user's interaction with the target interactive object includes: obtaining the user's second gaze duration on the target interactive object; and, if the second gaze duration reaches a second preset duration, determining that the user has performed an interaction with the target interactive object and obtaining the image enhancement data corresponding to the first illustration.

[0016] In this embodiment, users can manipulate the first illustration, the first sub-region within the first illustration, or the first main element simply by gazing at it, triggering image enhancement processing on the first illustration, the first sub-region within the first illustration, or the first main element. For users, this frees up their hands while still allowing them to perceive image enhancement, providing a better experience, especially in scenarios where it is inconvenient for users to operate the device.

[0017] In one optional embodiment, obtaining image enhancement data corresponding to the first illustration based on the user's interactive operation on the target interactive object includes: responding to the user's click operation on the target interactive object and obtaining image enhancement data corresponding to the first illustration.

[0018] In this embodiment, users can interact with the first illustration, the first sub-region of the first illustration, or the first main element by clicking, triggering image enhancement processing on the first illustration, the first sub-region of the first illustration, or the first main element. Especially when interactive prompts are displayed on the screen, users can intuitively experience the interaction process between themselves and the electronic device by clicking on the first illustration, the first sub-region of the first illustration, or the first main element according to the guidance of the interactive prompts, making it more interesting.

[0019] In an optional embodiment, obtaining image enhancement data corresponding to the first illustration based on the user's interactive operation on the target interactive object includes: responding to the user's interactive operation on the target interactive object through a touch preset interactive component and obtaining image enhancement data corresponding to the first illustration.

[0020] In this embodiment, interactive operations are performed on the first illustration, the first sub-region of the first illustration, or the first main element by touching the preset interactive components, which triggers image enhancement processing on the first illustration, the first sub-region of the first illustration, or the first main element. This method is applicable to most situations where one-handed operation of electronic devices is possible and has a wide range of applications.

[0021] In one optional embodiment, responding to the user's interaction operation on the target interactive object through a preset touch interaction component includes: responding to the user's touch operation on the preset interaction component based on a pre-registered touch monitoring event and determining the touch mode corresponding to the touch operation; and determining, through the first registration and callback interface bound to the touch monitoring event, the user's interaction operation on the target interactive object according to the touch mode.

[0022] In this embodiment, by providing diverse touch control methods, users can choose their preferred touch control method to operate preset interactive components. This is not only suitable for most situations where one hand is used to operate electronic devices, but also helps to enhance the fun of operation.

[0023] In an optional embodiment, the first interactive metadata includes display description information, which indicates the enhancement methods supported by the first illustration. Generating and displaying the first enhanced image based on the first interactive metadata and the image enhancement data includes: determining a first enhancement method supported by the first illustration based on the display description information; and generating and displaying the first enhanced image based on the first enhancement method and the image enhancement data. The first enhancement method can be a global enhancement method or a local enhancement method.

[0024] In this embodiment, the first interactive element indicates that the first illustration supports different enhancement methods. The first enhanced image corresponding to the first illustration can be generated according to the corresponding enhancement method, resulting in a variety of first enhanced images and enriching the display style of the first enhanced image.

[0025] In an optional embodiment, generating and displaying the first enhanced image based on the first enhancement method and the image enhancement data includes: if the first enhancement method is a global enhancement method, generating an enhanced image of the first illustration based on the image enhancement data as the first enhanced image; and displaying the first enhanced image on the display screen.

[0026] In an optional embodiment, generating and displaying the first enhanced image based on the first enhancement method and the image enhancement data includes: if the first enhancement method is a local enhancement method, generating an enhanced image of a first sub-region or a first main element in the first illustration based on the image enhancement data, as the first enhanced image; and displaying the first enhanced image on the display screen.

[0027] In this embodiment, based on whether the first illustration supports a global enhancement method or a local enhancement method, a first enhanced image corresponding to the corresponding enhancement method is generated and displayed. From the perspective of display style, not only is the purpose of enhancing the display effect of the first illustration achieved, but also the diverse display styles can bring a certain degree of fun to users, help stimulate users' interactive interest, and enhance the sense of immersion.

[0028] In one optional embodiment, obtaining the image enhancement data corresponding to the first illustration includes: obtaining the image enhancement data corresponding to the first illustration from a cloud server; or, obtaining the image enhancement data corresponding to the first illustration from a local preset storage space, wherein the image enhancement data is obtained in advance from the cloud server.

[0029] In this embodiment, pre-generating image enhancement data for the first illustration via a cloud server not only reduces the processing load on the electronic device, but also leverages the powerful capabilities of the cloud server to generate image enhancement data that better meets the image enhancement requirements, resulting in a first enhanced image with superior display quality and a better user experience. Furthermore, pre-retrieving the image enhancement data corresponding to the first illustration from the cloud server helps save time loading image enhancement data, resulting in higher processing efficiency and faster response times for the user, leading to a better overall experience.

[0030] In an optional embodiment, the display content processing method further includes: performing a sharpness enhancement process on the first illustration according to the first enhancement method to obtain a second enhanced image; and displaying the second enhanced image on the display screen.

[0031] In this embodiment, the first illustration is enhanced in sharpness by the electronic device itself, which can meet the user's image enhancement needs even in situations such as network outages or no network, thus providing greater flexibility.

[0032] In an optional embodiment, according to the first enhancement method, the first illustration is subjected to sharpness enhancement processing to obtain a second enhanced image, including: if the first enhancement method is a global enhancement method, the first illustration is subjected to sharpness enhancement processing to obtain a corresponding enhanced image, which is used as the second enhanced image.

[0033] In an optional embodiment, according to the first enhancement method, the first illustration is subjected to a sharpness enhancement process to obtain a second enhanced image, including: if the first enhancement method is a local enhancement method, a first sub-region or a first main element in the first illustration is subjected to a sharpness enhancement process to obtain a corresponding enhanced image, which is used as the second enhanced image.

[0034] In this embodiment, the electronic device can perform targeted clarity enhancement processing on the first illustration, the first sub-region in the first illustration, or the first main element according to the enhancement method supported by the first illustration, and display the second enhanced image corresponding to the first illustration, the first sub-region in the first illustration, or the first main element according to the corresponding enhancement method. The display method is more diversified, and the user experience is taken into account while enhancing the image display effect.

[0035] In one optional embodiment, obtaining the first illustration and the first interactive metadata includes: obtaining digital content from a cloud server, the digital content including structured text data, text metadata, at least one original image, and image metadata, wherein the at least one original image is used as at least one illustration in the text content, the at least one illustration includes the first illustration, and the text metadata or the image metadata includes the first interactive metadata; using a typesetting engine to typeset the text structured data and the at least one original image according to the text metadata and the image metadata to obtain corresponding typesetting information; and displaying the at least one illustration corresponding to the text content and the at least one original image on a display screen according to the typesetting information.

[0036] In this embodiment, by obtaining digital content containing the first illustration and the first interactive metadata from the cloud server, and by typesetting the corresponding digital content based on the typesetting engine, the corresponding text content and at least one illustration can be directly displayed on the screen. For electronic devices, there is no need to perform complex processing operations on the digital content, making it easier to use and more versatile.

[0037] In one alternative embodiment, displaying the first illustration on the display screen includes: determining a first display area occupied by the first illustration on the display screen based on the layout information; and displaying the first illustration within the first display area.

[0038] In this embodiment, the typesetting engine can be used to typeset multimodal digital content, adaptively display text content and at least one illustration on the display screen, and during the typesetting process, the display area of ​​each illustration on the display screen can be flexibly adjusted, making it suitable for displays with different display parameters and more versatile.

[0039] In one optional embodiment, obtaining the user's first gaze duration at the first illustration includes: determining the screen coordinates of the user's gaze point on the display screen; and obtaining the user's first gaze duration at the first illustration based on the screen coordinates and the first display area.

[0040] In one optional embodiment, determining the screen coordinates of the user's gaze point on the display screen includes: acquiring the user's eye information; determining the relative positional relationship between the user's gaze point and the display screen from the eye information based on pre-registered gaze point tracking events; and determining the screen coordinates of the user's gaze point on the display screen according to the relative positional relationship through a second registration and callback interface bound to the gaze point tracking events.

[0041] In this embodiment, based on the user's eye information, the relative position of the user's gaze point and the first illustration on the screen can be determined. Then, the duration of the user's gaze at the first illustration can be obtained to determine whether the user intends to gaze at the first illustration. This not only provides greater flexibility, but also ensures that the entire process is imperceptible to the user and does not impose any burden on the user.

[0042] In an optional embodiment, obtaining the user's first gaze duration at the first illustration based on the screen coordinates and the first display area includes: determining an inner transition area corresponding to the first display area based on the size of the first illustration and a first preset transition coefficient; determining whether the screen coordinates are located within the first display area based on the screen coordinates and the inner transition area; and if the screen coordinates are determined to be located within the first display area, obtaining the user's first gaze duration at the first illustration.

[0043] In an optional embodiment, the eye information includes first eye information obtained from the user within a preset sampling frequency, and the screen coordinates include a plurality of first screen coordinates determined based on the first eye information; determining whether the screen coordinates are located within the first display area based on the screen coordinates and the inner transition area includes: determining whether all of the plurality of first screen coordinates are located within a sub-display area excluding the inner transition area within the first display area; if it is determined that all of the plurality of first screen coordinates are located within the sub-display area, the screen coordinates are determined to be located within the first display area.

[0044] In an optional embodiment, the display content processing method further includes: determining an outer transition region corresponding to the first display area based on the size of the first illustration and a second preset transition coefficient; and determining that the screen coordinates are located outside the first display area if it is determined that the plurality of first screen coordinates are all located outside the outer transition region.

[0045] In this embodiment, corresponding inner and outer transition regions are set for the first display area according to the size of the first illustration and the first and second preset transition coefficients. The relative position relationship between the screen coordinates of the user's gaze point on the display screen and the first display area is continuously determined within a preset sampling frequency. This can determine whether the user's gaze point is stably falling on the first illustration, which helps to improve the accuracy of determining the user's gaze intention, avoids erroneous triggering of image enhancement processing functions, and helps to reduce the processing burden of electronic devices.

[0046] In an alternative embodiment, the first enhanced image is an image in which the main elements in the first illustration are displayed in a micro-motion manner.

[0047] In this embodiment, the main element in the first illustration is displayed in the first enhanced image by micro-motion, changing the main element from static display to dynamic display. This not only improves the display effect of the first illustration, but also enhances the realism and immersion for the user, thus improving the user experience.

[0048] In one alternative embodiment, the type of digital content includes digital picture books, digital catalogs, e-books, or digital games.

[0049] In this embodiment, digital content can take many forms, with a wider range of application scenarios. Different digital content can present different image enhancement effects during application, making it more interesting and offering richer display styles.

[0050] Secondly, embodiments of this application provide a data processing method for enhancing display content, applied to a first server. The method includes: acquiring at least one original image from digital content, the at least one original image being used as at least one illustration in text content, the digital content being used to display the text content and the at least one illustration on an electronic device; recognizing the at least one original image to identify element features of each main element from each of the at least one original image; and generating image enhancement data corresponding to each original image based on the element features of each main element in each original image.

[0051] In this embodiment, the first server performs recognition processing on each original image to determine the element features of the main elements included and generates corresponding image enhancement data that matches these element features. This processing method is simple and can be used to generate image enhancement data in batches, saving processing efficiency. Furthermore, the enhanced images generated based on the corresponding image enhancement data have a display effect that closely matches the illustration content, resulting in a more realistic feel and improving the user experience.

[0052] In an optional embodiment, the data processing method further includes sending the digital content and the image enhancement data to a second server.

[0053] In this embodiment, the first server, as the image enhancement data provider, provides image enhancement data to the second server, which is the user, thus helping to reduce the processing pressure on the user; or, the first server, as the master node server, provides image enhancement data to each of the second servers, which are edge servers. In a distributed server scenario, this helps each electronic device obtain image enhancement data from a nearby second server, thus expanding the application scope.

[0054] In an optional embodiment, the data processing method further includes: sending the digital content to the electronic device upon receiving a first acquisition request from the electronic device for the digital content.

[0055] In an optional embodiment, the data processing method further includes: upon receiving a second acquisition request from the electronic device for the first illustration, determining a first original image as the first illustration; acquiring image enhancement data corresponding to the first original image and sending it to the electronic device.

[0056] In this embodiment, the first server, as the user of image enhancement data or an edge server, directly communicates with the electronic device. In addition, the first server can provide corresponding data content for different acquisition requests from the electronic device, providing data support for the electronic device to perform image enhancement processing and display enhanced images.

[0057] In one optional embodiment, the at least one original image is identified to identify element features of each subject element from each of the at least one original image, including: inputting a first prompt word and the at least one original image into a first multimodal large model, instructing the first multimodal large model to identify the at least one original image to identify element features of each subject element from each of the at least one original image.

[0058] In this embodiment, based on the processing capabilities of the multimodal large model, the element features of each main element in each original image can be quickly and accurately identified. Furthermore, according to the characteristics of the multimodal large model, the recognition requirements and results can be adjusted at any time, resulting in greater flexibility.

[0059] In an optional embodiment, the data processing method further includes: determining that the interactive mode supported by the illustration is a global interactive mode, wherein the global interactive mode is used to indicate that the illustration corresponding to each original image supports interactive operations; and using the first coordinate range corresponding to each original image as interactive object information indicating the global interactive mode.

[0060] In this embodiment, when it is determined that the illustration supports the global interaction mode, the first coordinate range corresponding to each original image is used as the interaction object information indicating the global interaction mode. Based on this, the electronic device can determine that the illustration corresponding to each original image is an independent interaction object according to the first coordinate range, so that the user can perform a care operation on each illustration.

[0061] In an optional embodiment, the data processing method further includes: determining that the interactive mode supported by the illustration is a local interactive mode, wherein the local interactive mode is used to indicate that multiple sub-regions or multiple main elements in the illustration corresponding to each original image support interactive operations; and performing an interactive object marking operation on each original image according to the element characteristics of each main element in each original image, so as to mark the coordinate ranges corresponding to the multiple sub-regions or multiple main elements in each original image respectively.

[0062] In an optional embodiment, the data processing method further includes: determining a second coordinate range for marking multiple sub-regions or multiple main elements in each original image according to the interactive object marking operation; and using the multiple second coordinate ranges corresponding to each original image as interactive object information indicating the local interactive mode.

[0063] In this embodiment, when it is determined that the illustration supports a local interaction mode, an interactive object marking operation is performed on each original image. This marks the second coordinate range corresponding to multiple sub-regions or multiple main elements in each original image, and uses these multiple second coordinate ranges as interactive object information indicating the local interaction mode. Based on this, the electronic device can determine that multiple sub-regions or multiple main elements in each original image are independent interactive objects. In this way, users can interact with each illustration, and each sub-region or main element within an illustration, resulting in more diverse interaction methods and a better user experience.

[0064] In one optional embodiment, an interactive object labeling operation is performed on each original image based on the element features of each main element in each original image, including: inputting a second prompt word into the first multimodal large model, and instructing the first multimodal large model to perform an interactive object labeling operation on each original image based on the element features of each main element in each original image.

[0065] In this embodiment, based on the element features of each main element identified from each original image, the modal large model can be used to accurately and quickly perform interactive object marking operations on each original image. Especially for batch processing scenarios, the processing efficiency is higher, and the processing requirements and results can be adjusted at any time through prompt words, making it more flexible.

[0066] In an optional embodiment, the data processing method further includes: determining the enhancement methods supported by the illustration; generating display description information to indicate the enhancement methods; and generating interactive metadata corresponding to each original image based on the display description information and the interactive object information corresponding to each original image.

[0067] In this embodiment, the first server uses the display description information indicating the enhancement method and the interactive object information corresponding to each original image as the interactive metadata for each original image. Based on this, when the electronic device obtains the interactive metadata, it can determine the interactive mode and enhancement method supported by each illustration according to the parsing result of the interactive metadata. Then, the electronic device can display the enhanced image of each illustration in an adaptive manner to improve the display effect of the illustration.

[0068] In one optional embodiment, generating image enhancement data corresponding to each original image based on the element features of each main element in each original image includes: determining that the illustration supports a global enhancement method, inputting a third prompt word, each original image, and the element features of each main element in each original image into a second multimodal large model, and instructing the second multimodal large model to generate image enhancement data corresponding to each original image, wherein the global enhancement method is used to instruct the enhanced image of each illustration to be displayed separately.

[0069] In an optional embodiment, the data processing method further includes: upon receiving a second acquisition request from the electronic device for a first illustration among the at least one illustration, determining a first original image as the first illustration; acquiring image enhancement data corresponding to the first original image and sending it to the electronic device.

[0070] In an optional embodiment, image enhancement data corresponding to each original image is generated based on the element features of each main element in each original image, including: determining that the illustration supports a local enhancement method; inputting the fourth prompt word, each original image, the element features of each main element in each original image, and the marking result corresponding to the interactive object marking operation into a second multimodal large model; and instructing the second multimodal large model to generate image enhancement data corresponding to the sub-regions or main elements in each original image, wherein the local enhancement method is used to instruct the enhanced images of the sub-regions or main elements in each illustration to be displayed separately.

[0071] In an optional embodiment, the data processing method further includes: upon receiving a second acquisition request from the electronic device for a first illustration in the at least one illustration, determining a first original image as the first illustration; acquiring image enhancement data corresponding to multiple sub-regions or multiple main elements in the first original image and sending them to the electronic device.

[0072] In this embodiment, based on the enhancement methods supported by the illustration, the first server can generate corresponding image enhancement data individually for each original image, each sub-region within each original image, or each main element through a multimodal large model. Furthermore, upon receiving a second acquisition request from an electronic device for any interactive object among the first illustration, the first sub-region within the first illustration, or the first main element, the first server can send the corresponding image enhancement data to the electronic device. Based on this, after obtaining the corresponding image enhancement data, the electronic device can display the enhanced image corresponding to the interactive object, resulting in diversified display formats and enhancing both engagement and user experience.

[0073] In one optional embodiment, the digital content includes structured text data, text metadata, the at least one original image, and corresponding image metadata. The data processing method further includes: adding the interactive metadata corresponding to each original image to the corresponding image metadata; or, determining the association between each original image and the structured text data based on the text metadata and the image metadata; and adding the interactive metadata corresponding to each original image to the corresponding text metadata based on the association.

[0074] In an optional embodiment, the data processing method further includes: sending the digital content to the electronic device upon receiving a first acquisition request from the electronic device for the digital content.

[0075] In this embodiment, the interactive metadata corresponding to each original image is stored in text metadata or image metadata. Upon receiving a first retrieval request from an electronic device for the digital content, the digital content is sent to the electronic device. Based on this, the electronic device can simultaneously acquire the interactive metadata of each original image when acquiring the digital content, which helps to save on the number of interactions and improve processing efficiency.

[0076] Thirdly, embodiments of this application provide an electronic device, including: one or more processors; one or more memories; and one or more computer programs, wherein the one or more computer programs are stored in the one or more memories, and the one or more computer programs include instructions that, when executed by the one or more processors, cause the electronic device to perform the display content processing method as described in the first aspect.

[0077] Fourthly, embodiments of this application provide a server, including: one or more processors; one or more memories; and one or more computer programs, wherein the one or more computer programs are stored in the one or more memories, and the one or more computer programs include instructions that, when executed by the one or more processors, cause the server to perform the data processing method as described in the second aspect.

[0078] Fifthly, embodiments of this application provide a system comprising an electronic device and a server; the electronic device is configured to perform the display content processing method as described in the first aspect; and the server is configured to perform the data processing method as described in the second aspect.

[0079] Sixthly, embodiments of this application provide a chip that stores instructions, which, when executed, implement the data processing method as described in the second aspect, or implement the data processing method as described in the second aspect.

[0080] In a seventh aspect, embodiments of this application provide a computer-readable storage medium for storing a computer program that, when executed on an electronic device, causes the electronic device to implement the display content processing method as described in the first aspect or the data processing method as described in the second aspect.

[0081] Eighthly, embodiments of this application provide a computer program product, the computer program product including a computer program, which, when run on an electronic device, causes the electronic device to implement the display content processing method as described in the first aspect or the data processing method as described in the second aspect.

[0082] Based on the implementation methods provided above, the embodiments of this application can be further combined to provide more implementation methods.

[0083] It is understood that the solutions provided in the third to eighth aspects correspond to the solutions provided in the first and second aspects. Therefore, the beneficial effects of the third to eighth aspects can be referred to the beneficial effects of the first and second aspects, and repeated details will not be repeated. Attached Figure Description

[0084] Figure 1a This is a schematic diagram illustrating an interaction scenario between a cloud server and an electronic device, as provided in an embodiment of this application.

[0085] Figure 1b This is a schematic diagram illustrating another interaction scenario between a cloud server and an electronic device, provided as an embodiment of this application.

[0086] Figure 2 This is an example of an interaction signaling diagram between various functional modules of a cloud server provided in this application embodiment.

[0087] Figure 3 This is a structural block diagram of an electronic device provided in an embodiment of this application.

[0088] Figure 4a This is a signaling diagram illustrating the interaction between an electronic device and a cloud server, provided as an embodiment of this application.

[0089] Figure 4b This is a schematic diagram illustrating the relationship between the screen coordinates of a user's gaze point on the display screen and the relative position of a first display area on the display screen, as provided in an embodiment of this application.

[0090] Figure 4c This is a schematic diagram illustrating the relative positional relationship between the screen coordinates of a user's gaze point on the display screen and the transition area corresponding to the first display area on the display screen, as provided in an embodiment of this application.

[0091] Figure 5a This is a schematic diagram illustrating a user's interactive operation with the first illustration, provided as an embodiment of this application.

[0092] Figure 5b This is a schematic diagram illustrating a user's interactive operation on a first sub-region, provided as an embodiment of this application.

[0093] Figure 5c This is a schematic diagram illustrating another user interaction operation on the first sub-region provided in an embodiment of this application.

[0094] Figure 5d This is a schematic diagram illustrating another user interaction with the first illustration, provided as an embodiment of this application.

[0095] Figure 5e This is a schematic diagram illustrating another user interaction operation on the first sub-region provided in an embodiment of this application.

[0096] Figure 5f This is a schematic diagram illustrating another user interaction operation on the first sub-region provided in an embodiment of this application.

[0097] Figure 5g This is a schematic diagram illustrating another user interaction with the first illustration, provided as an embodiment of this application.

[0098] Figure 5h This is a schematic diagram illustrating another user interaction operation on the first sub-region provided in an embodiment of this application.

[0099] Figure 5i This is a schematic diagram illustrating another user interaction operation on the first sub-region provided in an embodiment of this application.

[0100] Figure 5j This is a schematic diagram illustrating the display of a first illustration on a screen of an electronic device, as provided in an embodiment of this application.

[0101] Figure 6 This is an interaction signaling diagram between various functional modules of an electronic device provided in an embodiment of this application.

[0102] Figure 7 A flowchart illustrating a data processing method for a cloud server, provided as an embodiment of this application.

[0103] Figure 8 This is a flowchart illustrating a method for processing display content in an electronic device, as provided in an embodiment of this application.

[0104] Figure 9a This is a structural block diagram of an electronic device provided in an embodiment of this application.

[0105] Figure 9b This is a software structure block diagram of an electronic device provided in an embodiment of this application.

[0106] Figure 10 This is a structural block diagram of another electronic device provided in an embodiment of this application. Detailed Implementation

[0107] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0108] It should be noted that, in the description of the embodiments of this application, unless otherwise stated, " / " means "or", for example, A / B can mean A or B. The term "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Furthermore, "at least one" refers to one or more, and "multiple" refers to two or more.

[0109] The term "exemplarily" means used as an example, embodiment, or illustration. Any embodiment illustrated herein as "exemplarily" is not necessarily to be construed as superior to or better than other embodiments. The terms "first," "second," etc., are used to distinguish similar objects and are merely a way of distinguishing objects with the same attributes in description. They do not limit the number or order of execution, and it should be understood that the words "first," "second," etc., do not necessarily imply that they are different.

[0110] Before detailing the technical solution of this application, it should be noted that while this embodiment uses a digital reading scenario as an example to illustrate the implementation principle, it is not limited to this in practical applications. For example, the technical solution of this application can also be applied to other fields such as games, photography, image restoration, medical image analysis, and industrial quality inspection. The specific application scenario can be selected according to actual needs. Furthermore, it should be noted that the type of e-book in the digital reading scenario is not limited. Optionally, it may include, but is not limited to, digital picture books, digital catalogs, e-books, or journals, etc., and the specific type can be determined according to the user's choice.

[0111] The application scenarios and implementation principles of the technical solution of this application will be explained below with reference to the accompanying drawings.

[0112] In traditional digital reading scenarios, illustrations are often added to ebooks to enhance the reading experience and reduce the difficulty of comprehension. However, simply relying on illustrations to aid reading is not ideal. Therefore, some manufacturers are introducing Augmented Reality (AR) or Virtual Reality (VR) technologies into digital reading, developing AR assets or VR scene data for ebooks to further enhance the reading experience. For ease of explanation, this AR / VR data will be referred to as image augmentation data. Based on this, by loading image augmentation data and displaying enhanced images corresponding to each illustration during the user's ebook reading process, the digital reading format can be enriched, and the reading experience can be enhanced.

[0113] However, the aforementioned image enhancement data not only requires a professional team and specialized tools to create, but the final generated image enhancement data is also quite large. Besides basic geometric data and system support data, it also includes a significant amount of dynamic interaction data, environmental perception data, and multimodal rendering data. Therefore, the entire production process is not only lengthy and costly, but electronic devices also experience considerable time consumption when downloading and loading the image enhancement data. Furthermore, loading the image enhancement data and rendering the corresponding enhanced images on electronic devices requires specific engines or hardware chips, making it difficult to guarantee universality in practical applications.

[0114] Therefore, embodiments of this application provide a data processing method and a display content processing method for enhancing display content, so as to reduce the difficulty and cost of producing image enhancement data and lower the display threshold of enhanced images, thereby improving the reading experience.

[0115] To clearly and completely explain the implementation principle of the technical solution of this application, the application process of the technical solution of this application in the digital reading scenario is described below.

[0116] In digital reading scenarios, device manufacturers provide reading cloud services for the electronic devices users use. These services offer a vast library of ebooks for users of various electronic devices. Correspondingly, device manufacturers also develop accompanying reading applications for the reading cloud service. These applications run on the electronic devices, allowing users to access the extensive ebook library and select their preferred titles. The reading cloud service is maintained by cloud servers, where the vast ebook resources are stored. Users can read online or download ebooks from the cloud server to their devices for offline reading, ensuring uninterrupted reading even without internet access and offering greater flexibility.

[0117] In this embodiment, the specific form of the cloud server is not limited. Optionally, the cloud server can be a single server, a distributed server, or a cluster server, and the specific form can be determined according to actual needs. Further optionally, regardless of the type of cloud server used by the device manufacturer, it can access a Content Delivery Network (CDN) to accelerate resource distribution, thereby saving time for electronic devices scattered in different geographical locations to download e-books.

[0118] Figure 1a This diagram illustrates an interaction scenario between a cloud server and an electronic device in a cloud reading service context.

[0119] like Figure 1a As shown in the diagram, the cloud server uses a distributed server as an example. The cloud server includes a master node server 10 and various edge servers 20 distributed across different geographical locations. Each edge server 20 has multiple electronic devices 30 within a preset range of its location, and each electronic device 30 can communicate with the edge server 20 within its preset range.

[0120] In this embodiment, the specific form of the plurality of electronic devices 30 is not limited; optionally, such as Figure 1a As shown, the multiple electronic devices 30 include, but are not limited to, smartphones, tablets, desktop computers, laptops, or e-readers, etc., and their specific forms can be determined according to actual needs.

[0121] like Figure 1a As shown in this embodiment, the master node server 10 is used to obtain the original data corresponding to newly uploaded e-books, and to parse and preprocess the corresponding original data to obtain compliant, valid, and standardized data. To distinguish it from the original data, in this embodiment, the data obtained by parsing and preprocessing the original data corresponding to the e-books is referred to as digital content.

[0122] Furthermore, such as Figure 1a As shown, after obtaining the digital content, the master node server 10 also distributes the corresponding digital content to each edge server 20. Based on this, each edge server 20, upon receiving a first acquisition request from each electronic device 30 within a preset range of its location, can send the digital content to the corresponding electronic device 30. Accordingly, as... Figure 1a As shown, after each electronic device 30 obtains the corresponding digital content from the edge server 20 within a preset range of its location, it can load and display the corresponding text content and illustrations locally for users to read.

[0123] In this embodiment, the preprocessing operations performed by the master node server 10 on the raw data of the e-book are not limited. Optionally, the preprocessing operations include, but are not limited to, format compliance detection, malicious code scanning, metadata standardization processing, security hardening processing, etc. In addition, in order to enhance the illustration effect displayed on each electronic device 30, the master node server 10 is also used to generate corresponding image enhancement data for each illustration in the e-book text content. After the corresponding image enhancement data is loaded on each electronic device 30, the enhanced image of the corresponding illustration can be displayed to enhance the display effect of the corresponding illustration.

[0124] Based on this Figure 1b This diagram illustrates another interaction scenario between a cloud server and an electronic device in a cloud service reading context.

[0125] like Figure 1b As shown, the master node server 10 can generate corresponding image enhancement data for each illustration in the e-book text content, and distribute the generated image enhancement data to each edge server 20. Based on this, as... Figure 1b As shown, each electronic device 30 can obtain image enhancement data from an edge server 20 within a preset range of its location. Thus, during user reading, when the electronic device 30 responds to the user's interactive operation on any illustration in the e-book text content—for ease of explanation, referred to as the first illustration—the electronic device 30 can display an enhanced image of the first illustration based on the obtained image enhancement data, thereby enhancing the display effect of the first illustration.

[0126] It should be noted that, Figure 1a and Figure 1b The illustration illustrates data communication between various electronic devices and an edge server. In practical applications, when the cloud server is a single, standalone server, it can both generate image enhancement data and communicate with various electronic devices. Regardless of the method used, the principle behind the cloud server's generation of image enhancement data remains the same, and will not be differentiated further in the following embodiments.

[0127] Next, in order to explain the technical solution of this application in detail, the terms, concepts and processing objects of cloud servers and electronic devices involved will be explained.

[0128] In this embodiment, the specific content of the e-book's raw data is not limited. Typically, the raw data of an e-book is multimodal data, optionally including, but not limited to, structured text data, text metadata, original images serving as illustrations within the text content, and image metadata. The structured text data is used to display the text content, and the text metadata indicates information such as the chapter, section, paragraph, and page number corresponding to the text content. Each original image serves as an illustration within the text content, and the image metadata indicates the attribute information of each original image, such as, but not limited to, the image identifier, storage location, format, size, and resolution of each original image.

[0129] Of course, the above information regarding various data is merely illustrative and is not limited to this in practical applications. The specific content can be determined based on the type of ebook, and will not be detailed here.

[0130] In this embodiment, the specific form of the original data is not limited. Optionally, the structured text data, text metadata, image metadata, and each original image can be separate and independent data objects. Alternatively, some of the data objects can be a whole. For example, image metadata can be embedded in the structured text data in binary, structured markup, or rich text formats. Regardless of the form, the association between the text content of the e-book and each illustration can be determined based on the text metadata and image metadata. For example, the text metadata can include the image identifiers corresponding to each illustration in the text content, and the image metadata can also indicate the chapter, section, paragraph, page number, etc., to which each original image belongs as an illustration in the text content.

[0131] In this embodiment of the application, the specific type of image enhancement data generated by the cloud server is not limited. Optionally, the image enhancement data can be an image file, such as a Web Picture or a Graphics Interchange Format (GIF), or a structured data object, such as an Extensible Markup Language (XML) or a Lightweight Data Interchange Format (JavaScript Object Notation, JSON), etc. The specific type can be selected according to actual needs, and will not be described in detail here.

[0132] In this embodiment, the type of enhanced image displayed on the electronic device is not limited. Optionally, the type of enhanced image includes, but is not limited to, animated GIFs, videos, high-definition still images, etc., and the type of enhanced image can vary depending on the actual needs. For example, for e-book types with high interactivity, such as children's picture books, the enhanced image corresponding to the illustrations can be an animated GIF or a video. For e-book types such as books, journals, and popular science books, the enhanced image corresponding to the illustrations can be a high-definition still image. Of course, this is only an illustrative example.

[0133] In this embodiment, while reading an e-book, a user can interact with the first illustration in the text content to trigger the electronic device to display an enhanced image of the first illustration. The type of interaction is not limited; optionally, this embodiment pre-sets multiple interaction types, including but not limited to user interaction with the first illustration through gazing, clicking, or touch operations on preset interactive components. Optionally, the user can use touch methods such as tapping, long tapping, swiping, rotating, and continuous tapping to operate the preset interactive components. The specific touch method can be determined according to actual needs and is not limited here.

[0134] In this application embodiment, the specific timing of the electronic device obtaining image enhancement data from the cloud server is not limited. Depending on the actual situation of the electronic device, the timing of obtaining image enhancement data from the cloud server may also be different. Several optional methods are illustrated below.

[0135] In one alternative approach, if the electronic device has sufficient storage space, it can simultaneously acquire all the image enhancement data corresponding to the digital content of the ebook from the cloud server. This way, when the user interacts with the first illustration while reading the ebook, the electronic device can load the image enhancement data corresponding to the first illustration locally and display the enhanced image of the first illustration, resulting in faster loading and a better user experience.

[0136] In another alternative approach, to balance storage space and loading speed, the electronic device first retrieves the digital content of the ebook from a cloud server and loads and displays the text content and illustrations locally. Based on this, during the user's reading process, the electronic device synchronously retrieves image enhancement data corresponding to each illustration in the relevant text content from the cloud server according to the user's current reading progress. For example, during reading, the electronic device identifies the chapter, section, paragraph, and page number corresponding to the currently read text content, determines all illustrations associated with the corresponding chapter, section, paragraph, and page number, and retrieves the image enhancement data corresponding to all illustrations from the cloud server. Furthermore, in response to the user's interactive operation on the first illustration, the electronic device loads the pre-retrieved image enhancement data for the first illustration locally and displays the enhanced image of the first illustration. In this way, the electronic device does not need to retrieve the image enhancement data for all illustrations in the ebook at once, saving storage space while ensuring loading speed.

[0137] In another alternative approach, if the network is stable and has high bandwidth, the electronic device can retrieve the digital content of the ebook from a cloud server, load the corresponding digital content locally, and display the text content and illustrations of the ebook. Based on this, during the user's reading process, when the electronic device responds to the user's interactive operation on the first illustration, it directly retrieves the image enhancement data corresponding to the first illustration from the cloud server. Furthermore, upon retrieving the corresponding image enhancement data, it loads the corresponding image enhancement data and displays the enhanced image of the first illustration, thus maximizing storage space savings.

[0138] It should be noted that the above methods are only illustrative examples. In actual applications, they are not limited to these methods. The specific method to be used can be selected according to actual needs.

[0139] In this embodiment, the illustrations in the e-book text content support different interaction modes, including global interaction mode and local interaction mode. This can be understood as follows: for illustrations supporting global interaction mode, the illustration itself acts as an independent interactive object, allowing the user to interact with the entire illustration; for illustrations supporting local interaction mode, sub-regions or main elements within the illustration act as independent interactive objects, allowing the user to interact with these sub-regions or main elements. Here, a sub-region in the illustration refers to the original image being divided into multiple sub-regions, each of which serves as a sub-region of the illustration. Different main elements in the illustration can be understood as the various contents displayed within the illustration. For example, in an illustration displaying mountains, water, clouds, trees, and birds, the mountains, water, clouds, trees, and birds are each different main elements.

[0140] In this embodiment, the illustrations in the e-book text content also support different enhancement methods, including global enhancement and local enhancement. This can be understood as follows: regardless of whether the illustration supports a global or local interaction mode, after the user interacts with the illustration or its sub-regions or main elements, the enhanced image of the entire illustration is displayed; this is a global enhancement method. For illustrations supporting local interaction mode, after the user interacts with a sub-region or main element of the illustration, the enhanced image of the corresponding sub-region or main element is displayed; this is a local enhancement method.

[0141] In this application embodiment, the setting method for the interactive mode and enhancement method supported by the illustration is not limited. Optionally, the provider can pre-set the method in the original data before the e-book is newly listed, or the cloud server can dynamically set the method in the process of preprocessing the original data or generating image enhancement data after receiving the original data corresponding to the newly listed e-book.

[0142] Furthermore, when setting the interaction modes and enhancement methods supported by illustrations, you can set them from the perspective of the entire ebook or from the perspective of individual illustrations. For example, if you set them from the perspective of the entire ebook, all illustrations in the ebook's text content will support the same interaction modes and enhancement methods; as another example, if you set them from the perspective of individual illustrations, different illustrations in the ebook's text content can support different interaction modes and enhancement methods.

[0143] Alternatively, taking the setting from the perspective of the entire ebook as an example, the following is an exemplary description of the process of dynamically setting the interactive modes and enhancements supported by illustrations on the cloud server.

[0144] In one alternative approach, during the preprocessing of the ebook's raw data or the generation of image enhancement data, the cloud server can identify individual original images. If most of the identified original images have complex content—for example, the main elements in the display content have blurred or no clear boundaries, or different main elements are intertwined and blended—the cloud server can set the interactive mode supported by the illustrations to a global interactive mode and the enhancement method supported by the illustrations to a global enhancement method from the perspective of the entire ebook.

[0145] In another alternative approach, during the preprocessing of the ebook's raw data or the generation of image enhancement data, the cloud server can identify individual original images. If most of the identified original images have complex content—for example, the main elements have clear boundaries and different main elements are independent of each other—the cloud server can set the interactive modes supported by the illustrations to local interactive modes from the perspective of the entire ebook.

[0146] Optionally, the cloud server can flexibly configure the illustration enhancement methods based on the type of ebook. For example, for ebooks with high interactivity requirements, such as children's picture books, the cloud server can set the illustration enhancement method to local enhancement to enhance reading enjoyment. This way, when children interact with illustrations in the text by clicking, they can achieve a "click-and-action" effect, which helps to motivate their reading. Conversely, for ebooks with relatively low interactivity requirements, such as periodicals and popular science books, the cloud server can set the illustration enhancement method to global enhancement. This way, when users interact with illustrations in the text, the entire illustration's enhanced image is displayed, improving the illustration's display effect and enhancing the reading experience.

[0147] It should be noted that the above examples are intended to illustrate the dynamic configuration principle of cloud servers and are not limiting. In practical applications, the specific configuration method can be determined according to actual needs, and will not be detailed here.

[0148] For ease of explanation, in this embodiment, the interactive modes and enhancement methods supported by the illustrations are described by way of the provider's pre-setting from the perspective of the entire e-book, and will not be repeated in subsequent embodiments.

[0149] In this embodiment, based on the different types of subject elements, the subject elements included in each original image are distinguished into dynamic subject elements and static subject elements. Dynamic subject elements refer to subject elements with dynamic characteristics in a real-world scene, while static subject elements refer to subject elements with static characteristics in a real-world scene. For example, subject elements such as clouds, seawater, and flags have dynamic characteristics in a real-world scene; if the original image includes these subject elements, it is determined to be a dynamic subject element. Conversely, subject elements such as land, mountains, and buildings have static characteristics in a real-world scene; if the original image includes these subject elements, it is determined to be a static subject element.

[0150] Based on this, when generating corresponding image enhancement data from each original image, the cloud server can generate image enhancement data that adapts to the features of each subject element in each original image. In this way, when electronic devices generate enhanced images of each illustration based on the corresponding image enhancement data, the display effect is more realistic.

[0151] Based on the above, the process of generating image enhancement data on the cloud server will be explained in detail below.

[0152] In this embodiment of the application, the specific functional modules included in the cloud server are not limited. Optionally, the cloud server may include a parsing module, a preprocessing module, an identification module, and a generation module.

[0153] Based on this Figure 2 The diagram illustrates the signaling interactions between various functional modules of the cloud server during the image enhancement data generation process.

[0154] like Figure 2 As shown in steps (1) and (2), the parsing module is used to receive the original data corresponding to the newly uploaded e-book and to parse the corresponding original data to obtain the text structured data, text metadata, original images of each illustration in the text content of the e-book, and image metadata. For the specific content of each data item, please refer to the explanation of the foregoing embodiment, which will not be repeated here.

[0155] Furthermore, after parsing the above data, the parsing module, such as Figure 2 As shown in steps (3) and (4), the parsing module is also used to input the above data into the preprocessing module. The preprocessing module is used to perform preprocessing operations on the above data to obtain compliant, valid, and standardized data. Optionally, the preprocessing operations include, but are not limited to, performing format compliance checks, malicious code scanning, metadata standardization processing, and security hardening processing on the data. The specific processing operations can be selected according to actual needs and are not limited here.

[0156] Furthermore, in order to generate image enhancement data, such as Figure 2 As shown in step (5), the preprocessing module is also used to obtain each original image. For example, the preprocessing module can determine the storage location of each original image based on the image metadata, and then obtain the corresponding original image based on the storage location of each original image.

[0157] Furthermore, after obtaining each original image, the preprocessing module, such as... Figure 2As shown in steps (6) and (7), the preprocessing module is also used to input each original image into the recognition module. The recognition module is used to perform recognition processing on each original image to identify the element features of each main element in each original image. Among them, the element features corresponding to each main element include, but are not limited to, the type, color, texture, shape outline and boundary information of each main element.

[0158] Furthermore, after the recognition module identifies the element features of each main element from each original image, such as... Figure 2 As shown in steps (8) and (9), element features for each main element in each original image are identified and returned to the preprocessing module. Then, the preprocessing module returns the data to the parsing module. Based on this, after receiving the element features of each main element in each original image returned by the preprocessing module, as shown in steps (8) and (9), the parsing module returns the data to the parsing module. Figure 2 As shown in step (10), the parsing module is used to add the element features of each subject element to the text metadata or image metadata.

[0159] In this embodiment of the application, in order to determine whether each illustration in the e-book text content is an independent interactive object, such as Figure 2 As shown in step (11), the recognition module is also used to determine the pre-set interaction mode for illustrations in the e-book text content, in order to determine whether to perform interactive object marking operations on each original image. Based on this, as Figure 2 As shown in step (12), if the recognition module determines that the pre-set interaction mode is a local interaction mode, it performs an interaction object marking operation on each original image so that the corresponding image enhancement data can be generated based on the marking results. The specific method of the interaction object marking operation is not limited. Optionally, the recognition module can divide each original image into multiple sub-regions according to a preset division rule and mark the coordinate range of each sub-region, where each sub-region serves as an independent interaction object in the corresponding illustration; or, the recognition module can determine multiple main elements in each original image based on the element features identified from each original image and mark the coordinate range of each main element, where each main element serves as an independent interaction object in the corresponding illustration.

[0160] Furthermore, such as Figure 2 As shown in steps (13) and (14), the recognition module is also used to return the coordinate range corresponding to each sub-region or main element in each original image to the preprocessing module, and then the preprocessing module returns the above data to the parsing module. Based on this, after receiving the coordinate range corresponding to each sub-region or main element in each original image, the parsing module, as shown in steps (13) and (14), returns the data to the parsing module. Figure 2As shown in step (15), the parsing module is also used to add the coordinate ranges corresponding to each sub-region or main element in each original image to the text metadata or image metadata as interactive metadata, so as to indicate that the illustrations corresponding to each original image support local interactive mode. It can be understood that if the text metadata or image metadata corresponding to each original image only includes the coordinate range corresponding to the corresponding original image, it means that the illustrations corresponding to the corresponding original image support global interactive mode; if it includes the coordinate ranges corresponding to each sub-region or main element in the corresponding original image, it means that the illustrations corresponding to the corresponding original image support local interactive mode.

[0161] In addition, after the preprocessing module receives the element features of each main element in each original image returned by the recognition module, such as... Figure 2 As shown in step (16), the preprocessing module is also used to input the element features of each original image and its corresponding main elements into the generation module so that the generation module can generate image enhancement data based on the above data.

[0162] Based on this, after the generation module receives the original images and their corresponding element features from the preprocessing module, such as... Figure 2 As shown in step (17), the generation module is used to determine the image enhancement parameters corresponding to each subject element based on the element features of each subject element in each original image, so as to generate the corresponding image enhancement data in the future. The specific content of the image enhancement parameters is not limited. Optionally, the image enhancement parameters corresponding to each subject element include, but are not limited to, parameters such as sharpness, color, transparency, saturation, brightness, and shadow. For dynamic subject elements, they may also include parameters such as the dynamic effect, motion trajectory, duration, frame rate, and playback mode corresponding to the subject element. The specific parameter content can be determined according to the subject element type, and will not be detailed here.

[0163] Furthermore, such as Figure 2 As shown in step (18), before generating image enhancement data, the generation module is also used to determine the enhancement method pre-set for illustrations in the e-book text content, so as to generate image enhancement data in different ways in the future.

[0164] Based on this, if the generation module determines that the illustration supports a global enhancement method, such as Figure 2 As shown in steps (19)-(21), the generation module can generate image enhancement data corresponding to each original image based on the element features of each main element in each original image and the image enhancement parameters, and return the image enhancement data corresponding to each original image to the preprocessing module, and then return it to the parsing module through the preprocessing module.

[0165] Accordingly, if the generation module determines that the enhancement method supported by the illustration is a local enhancement method, such as Figure 2 As shown in steps (22)-(24), the generation module can generate image enhancement data corresponding to each sub-region or main element in each original image based on the marking results obtained by performing interactive object marking operations on each original image, and return the image enhancement data corresponding to each sub-region or main element in each original image to the preprocessing module, and then return it to the parsing module through the preprocessing module.

[0166] Optionally, after determining the enhancement methods supported by the illustration, the generation module can also generate display description information to indicate the corresponding enhancement methods. Figure 2 (Not shown in the image). Based on this, when the generation module returns image enhancement data to the preprocessing module and the parsing module, it can also synchronously return the corresponding display description information, so that the parsing module can add the corresponding display description information to the text metadata or image metadata corresponding to each original image as interactive metadata.

[0167] It should be noted that, Figure 2 The steps shown are merely illustrative examples intended to illustrate the implementation principle of the display content processing method provided in this application embodiment by the cloud server. In practical applications, the specific processing steps are not limited to these and will not be elaborated here.

[0168] Based on the above, after the cloud server generates image enhancement data for each original image through the cooperation of the various functional modules, it can use the CDN network for acceleration and resource distribution so that various electronic devices can download it on demand.

[0169] In this embodiment, the specific methods by which the recognition module performs recognition processing on each original image and the generation module generates image enhancement data are not limited; refer to... Figure 2 The following is an exemplary description of the processing procedures of the identification module and the generation module.

[0170] In this embodiment, the recognition module can perform recognition processing on each original image using traditional image processing algorithms or image processing tools, or it can perform recognition processing on each original image using a multimodal large language model (MLLM) with image recognition capabilities. The specific method can be selected according to actual needs.

[0171] Optionally, the following explanation uses the example of the recognition module performing recognition processing on each original image through the first multimodal large model.

[0172] In this embodiment of the application, the recognition module may pre-set a first prompt word template, which includes a first prompt word used to instruct the first multimodal large model to perform recognition processing on each input original image in order to identify the element features of each main element from each original image.

[0173] Alternatively, the following shows an alternative form of the first prompt word template:

[0175] You are a professional image recognition assistant, please perform the following operations:

[0176] 1. First, provide a global description for each image (within 50 characters);

[0177] 2. Identify the element features of each main element in each image, such as, but not limited to, type, color, texture, and shape outline;

[0178] 3. By default, it prioritizes recognizing dynamic objects in the foreground of each image;

[0179] 4. Output the identified data in a structured format. 】

[0181] Based on the above, the recognition module can input each original image and the first prompt word template into the first multimodal large model, instructing the first multimodal large model to identify the element features of each main element from each original image. Furthermore, after identifying the element features of each main element from each original image through the first multimodal large model, the recognition module can also determine the interaction mode pre-set by the provider for the illustrations in the ebook text content, in order to determine whether interactive object tagging operations need to be performed on each original image.

[0182] In one optional embodiment, the recognition module determines that the interaction mode pre-set by the provider is a global interaction mode. In this case, the recognition module can directly return the element features of each subject element in each original image recognized by the first multimodal method to the preprocessing module, so that the preprocessing module can input each original image and the element features of each subject element in each original image into the generation module, and instruct the generation module to generate image enhancement data corresponding to each original image.

[0183] In another embodiment, the identification module determines that the interaction mode pre-set by the provider is a partial interaction mode. (See also...) Figure 2 In this case, in addition to returning the element features of each main element in each identified original image to the preprocessing module, the recognition module is also used to perform interactive object labeling operations on each original image so that the generation module can generate corresponding image enhancement data based on the labeling results of each original image.

[0184] The method by which the recognition module performs interactive object labeling operations on each original image is not limited. Optionally, while performing image recognition on each original image using the first multimodal large model, the recognition module can also perform interactive object labeling operations on each original image using the first multimodal large model. Based on this, the first prompt word template can also include a second prompt word to instruct the first multimodal large model to perform interactive object labeling operations on each original image.

[0185] Alternatively, another alternative form of the first prompt word template is shown below:

[0187] You are a professional image recognition assistant, please perform the following operations:

[0188] 1. First, provide a global description for each image (within 50 characters);

[0189] 2. Perform the following operations on each image:

[0190] 1) Identify the element features of each main element in each image, such as, but not limited to, type, color, texture, and shape outline;

[0191] 2) Based on the element characteristics of each main element, determine the outline of each main element and mark it with boundary lines;

[0192] 3) Mark the coordinate range of the movable and static main elements in each image, in the format: [x1,y1,x2,y2];

[0193] 4) Determine the segmentation mask data corresponding to each main element in each image;

[0194] 3. By default, it prioritizes recognizing dynamic objects in the foreground of each image;

[0195] 4. Output the identified data in a structured format. 】

[0197] Based on this, the first multimodal large model, using the first cue word template, can identify the element features of each main element in each original image. Then, based on the identified element features, it determines the contour and corresponding coordinate range of each main element in each original image, as well as identifies dynamic main elements with dynamic features and static main elements with static features. Based on this, it marks the contour, coordinate range, and information such as movable and static main elements in each original image.

[0198] Optionally, when determining the main elements that can be used as independent interactive objects, the marked movable main elements can be treated as independent interactive objects. For example, an original image includes main elements such as "mountain, water, cloud, tree, and bird." The first multimodal large model identifies the element features corresponding to each of these main elements, determining that "water, cloud, and bird" have dynamic features, while "mountain and tree" have static features. Based on this, according to the identified element features of each main element, the outline and coordinate range of each main element are marked, and "mountain and tree" are marked as static areas, while "water, cloud, and bird" are marked as movable main objects that can be used as independent interactive objects.

[0199] It should be noted that in the above example, the method of marking interactive objects is illustrated by identifying the main elements of each original image and marking the coordinate range of each main element. In practical applications, the second prompt word can also specify dividing each original image into regions and marking the coordinate range of each sub-region. The implementation process is similar to the above example, and can be found in the aforementioned explanation, which will not be repeated here.

[0200] Based on this, the recognition module can identify the element features of each main element in each original image and the coordinate range corresponding to each sub-region or main element in each original image according to the recognition results of the first multimodal large model. Then, the element features of each main element in each original image and the coordinate range corresponding to each sub-region or main element are returned to the preprocessing module. The preprocessing module inputs the above data into the generation module, instructing the generation module to generate image enhancement data corresponding to each sub-region or main element in each original image.

[0201] Furthermore, to generate image enhancement data, after receiving the element features of each subject element in each original image from the preprocessing module, and the coordinate range corresponding to each sub-region or subject element in each original image, the generation module can determine the image enhancement parameters for each subject element based on its element features. Then, based on the image enhancement parameters for each subject element in each original image, image enhancement data corresponding to each original image and each sub-region or subject element in each original image is generated.

[0202] In this embodiment, the specific method by which the generation module determines the image enhancement parameters of each subject element is not limited. Optionally, the generation module can also determine the image enhancement parameters of each subject element through a multimodal large model.

[0203] In this embodiment, the example given is that the generation module determines the image enhancement parameters of each subject element using a second multimodal large model. Optionally, the generation module presets a second prompt word template, including a fifth prompt word, to instruct the second multimodal large model to determine the image enhancement parameters of each subject element based on each input original image and the element features of each subject element in each original image.

[0204] Alternatively, for a second cue word template that includes a fifth cue word, an alternative form of the second cue word template is shown below:

[0205]

[0206]

[0207] It should be noted that the specific content of the fifth prompt word is not limited in the embodiments of this application. In the above example, the example of instructing the second multimodal large model to determine multiple image enhancement parameters by using a fifth prompt word is used for illustration. In practical applications, the generation module can also preset multiple fifth prompt words to instruct the second multimodal large model to determine different image enhancement data respectively.

[0208] For example, a fifth prompt word can be used to instruct the second multimodal large model to determine the basic motion effect attributes of the movable area. The corresponding prompt word content can be information such as instructing the generation of a mask for the dynamic subject element and defining the range of motion effect. Another example is a fifth prompt word used to instruct the second multimodal large model to determine features such as motion direction and deformation probability. The corresponding prompt word content can be information such as instructing the definition of motion effect types (fade-in / fade-out, translation, scaling, rotation, etc.) and parameters (duration, frame rate, etc.) for the dynamic subject element. Yet another example is a fifth prompt word used to instruct the second multimodal large model to determine the scene attribute parameters of the dynamic subject element. The corresponding content can be information such as instructing the definition of attribute information (playback order, delay time, and area mapping relationship) for the dynamic subject element.

[0209] Based on this, the generation module can sequentially input different fifth cue words into the second multimodal large model to instruct the second multimodal large model to determine different types of image enhancement parameters corresponding to each subject element in each original image according to each fifth cue word.

[0210] Optionally, after determining the image enhancement parameters for each subject element using the second multimodal large model, the generation module can also synchronously generate corresponding image enhancement data based on the elemental features of each subject element in each original image and the image enhancement parameters determined above. In this way, the determination of image enhancement parameters and the generation of image enhancement data are completed in one go using the second multimodal large model, resulting in low error and high efficiency throughout the entire process.

[0211] Reference Figure 2 In order to determine the method for generating image enhancement data, the generation module is also used to determine the enhancement method pre-set for illustrations in the text content, so as to generate corresponding image enhancement data according to the different enhancement methods supported by the illustrations.

[0212] In an optional embodiment, the generation module determines that the pre-set enhancement method for illustrations in the text content is a global enhancement method. In this case, the second prompt word template may also include a third prompt word, which instructs the second multimodal large model to generate image enhancement data corresponding to each original image based on the element features of each subject element in each original image and the image enhancement parameters.

[0213] Alternatively, for a second cue word template that includes a third cue word, an alternative form of the second cue word template is shown below:

[0214]

[0215]

[0216] Based on this, the generation module inputs the second prompt word template, each original image, and the element features of each main element in each original image into the second multimodal large model. The second multimodal large model can then determine the image enhancement parameters of each main element based on the second prompt word template, and generate image enhancement data corresponding to each original image based on the element features and image enhancement parameters of each main element in each original image. This can be understood as the JSON format configuration file generated for each original image being the corresponding image enhancement data.

[0217] Furthermore, after obtaining the image enhancement data corresponding to each original image, the generation module can also establish a first mapping relationship between each original image and the corresponding image enhancement data. For example, it can associate each original image with the identifier of the corresponding image enhancement data so that when the electronic device initiates a second acquisition request for the first illustration, the image enhancement data corresponding to the first original image as the first illustration can be obtained according to the first mapping relationship.

[0218] In another optional embodiment, the generation module determines that the pre-set enhancement method for the illustrations in the text content is a local enhancement method. In this case, the second prompt word template may also include a fourth prompt word, which instructs the second multimodal large model to generate image enhancement data corresponding to each sub-region or main element based on the element features and image enhancement parameters corresponding to each sub-region or main element in each original image.

[0219] Alternatively, for a second cue word template that includes a fourth cue word, another alternative form of the second cue word template is shown below:

[0220]

[0221]

[0222] Based on this, the generation module can input the second prompt word template containing the fourth prompt word, each original image, and the element features of each subject element in each original image into the second multimodal large model. This instructs the second multimodal large model to generate image enhancement data corresponding to each subject element in each original image based on the element features and image enhancement parameters of each subject element in each original image after determining the image enhancement parameters of each subject element.

[0223] This can be understood as follows: the JSON object generated for each main element in each original image is the image enhancement data corresponding to the corresponding main element; and the image enhancement data corresponding to each main element in each original image is the image enhancement data corresponding to the original image.

[0224] It should be noted that in the above example, the image enhancement data corresponding to each subject element in each original image is generated by the second multimodal large model. In practical applications, the fourth prompt word can also instruct the generation of corresponding image enhancement data for each sub-region in each original image. The implementation process is similar to the above example, and can be found in the foregoing explanation, which will not be repeated here.

[0225] Furthermore, after obtaining the image enhancement data corresponding to each sub-region or main element in each original image, the generation module can also establish a second mapping relationship between each sub-region or main element in each original image and the corresponding image enhancement data. This allows the electronic device to obtain the image enhancement data corresponding to the first sub-region or first main element according to the second mapping relationship when the electronic device initiates a second acquisition request for the first sub-region or first main element in the first illustration. The first sub-region corresponds to the first sub-region in the first illustration.

[0226] Optionally, after performing interactive object marking operations on each original image, the recognition module can also establish corresponding object identifiers for each sub-region or main element in each original image. Based on this, the generation module can associate the object identifiers corresponding to each sub-region or main element in each original image with the identifiers of the corresponding image enhancement data as a second mapping relationship.

[0227] In the above embodiments, the following are examples: each original image is identified and processed by the first multimodal large model, and the image enhancement parameters of each subject element are determined by the second multimodal large model and image enhancement data corresponding to each original image is generated, or image enhancement data corresponding to each sub-region or subject element in each original image is generated.

[0228] In practical applications, all of the above operations can also be achieved through a general multimodal large model. Accordingly, the prompt word template used to instruct the general multimodal large model to perform all of the above operations can be a general prompt word template containing multiple prompt words. For example, the general prompt word template can simultaneously include the first prompt word, the second prompt word, the third prompt word, the fourth prompt word, and the fifth prompt word.

[0229] Alternatively, for illustrations that support both local interactive modes and global enhancements, an optional form of the general prompt word template is shown below:

[0230]

[0231]

[0232] Based on this, by inputting the above-mentioned general prompt word template and each original image into the general multimodal large model, the general multimodal large model can be instructed to complete the identification of the main elements of each original image, mark the outline and coordinate range of each main element in each original image, determine the image enhancement parameters of each main element, and generate the image enhancement data corresponding to each original image in one go. The whole process is uninterrupted and has a lower error rate.

[0233] It should be noted that the various prompt word templates and their contents mentioned in the above embodiments are merely illustrative examples. In actual applications, they are not limited to these and can be flexibly adjusted according to actual needs. They will not be detailed here.

[0234] Alternatively, to improve the accuracy of image enhancement data, before determining the image enhancement parameters of each subject element in each original image and generating image enhancement data through the second multimodal large model, the cloud server can also optimize the processing process or results of the second multimodal large model based on the semantic relevance between the image content of each original image and the text content of the e-book, so that the enhanced image displayed on the electronic device is more compatible with the text content of the e-book and the user experience is better.

[0235] In one optional approach, the cloud server pre-sets a third prompt word template. This template instructs the third multimodal large-scale model to determine the preceding and following text paragraphs associated with each original image and its associated text paragraphs, perform semantic analysis on each original image and its associated text paragraphs, and optimize the prompt word content of the second prompt word template based on the analysis results. Based on this, the optimized second prompt word template guides the second multimodal large-scale model to redetermine image enhancement parameters and generate image enhancement data, thereby improving the accuracy of both the image enhancement parameters and the image enhancement data.

[0236] In another alternative approach, the cloud server pre-sets a fourth prompt word template, which is used to instruct the fourth multimodal big model to determine the preceding and following text paragraphs associated with each original image based on each input original image and text content, perform semantic analysis on each original image and its associated preceding and following text paragraphs, and optimize the image enhancement data generated by the second multimodal big model based on the analysis results, thereby improving the accuracy of the corresponding image enhancement data.

[0237] Regardless of the method used, in the above embodiments, the optimization process based on semantic analysis results can be an iterative process. Thus, through continuous iterative processing, until the error value of the image enhancement data is less than a preset error threshold, an enhanced image with satisfactory display results is obtained.

[0238] In this embodiment, the form and content of the third and fourth prompt word templates are not limited. Optionally, refer to the exemplary descriptions of the first and second prompt word templates in the foregoing embodiments, which will not be repeated here.

[0239] It should be noted that the specific types of the first multimodal large model, the second multimodal large model, the third multimodal large model, and the fourth multimodal large model mentioned in the above embodiments are not limited. The model type selected may vary depending on the actual needs.

[0240] It should be further noted that the types of various multimodal models can be the same or different. The terms "first" and "second" are only used to distinguish the different functions implemented by the multimodal models, and do not distinguish the types of multimodal models.

[0241] In this application, the cloud server can obtain image enhancement data that highly matches the element features of each main element in each original image by performing the above processing operations. In this way, after the electronic device loads each image enhancement data, it can display the enhanced image corresponding to the content of the illustration. The dynamic main elements in the illustration are displayed with dynamic effects in the displayed enhanced image, and the static main elements are displayed with better static effects in the displayed enhanced image.

[0242] For example, if the dynamic subject element in the illustration is "seawater," then the enhanced image will display "seawater" with a ripple effect; if the dynamic subject element in the illustration is "flag," then the enhanced image will display "flag" with a fluttering effect; if the dynamic subject element in the illustration is "bird," then the enhanced image will display "bird" with a flying effect; if the static subject element in the illustration is "mountains and rivers," then the enhanced image will display "mountains and rivers" in high definition, high saturation, or more vivid colors, and so on.

[0243] Based on this, by displaying enhanced images corresponding to various illustrations in the text content, electronic devices not only allow users to understand the text content more clearly and intuitively, reducing the difficulty of comprehension, but also enrich the form of digital reading and provide a better user experience.

[0244] In the above embodiments, the process of generating enhanced image data for each illustration in the text content of an e-book by a cloud server has been described in detail. Below, the electronic device provided in the embodiments of this application and its specific functions will be described in detail with reference to the accompanying drawings.

[0245] Figure 3 A schematic diagram of the structure of an electronic device 300 is shown, such as Figure 3 As shown in this embodiment, the electronic device 300 includes a front-facing camera 310 and a first processing chip 320 associated with the front-facing camera 301. The front-facing camera 310 continuously collects eye information from the user while using the electronic device 300, and the first processing chip 320 identifies and processes the eye information collected by the front-facing camera 310 to determine the relative position of the user's gaze point and the display screen.

[0246] In this application embodiment, the specific type of the front camera 310 is not limited. Optionally, in order to save processing resources, the front camera 310 can be an independent low-power camera. When the electronic device 300 is in standby mode, the front camera 310 is always on to collect the user's eye information at any time.

[0247] Accordingly, the specific structure of the first processing chip 320 is not limited. Optionally, the first processing chip 320 includes a preset algorithm model and a neural network processing unit (NPU) or a graphics processing unit (GPU). The preset algorithm model is used in conjunction with the NPU or GPU to perform operations such as parsing, recognizing, denoising, and feature extraction on the eye information collected by the front-facing camera 310, as well as to identify the user's gaze point from the corresponding eye information and determine the relative positional relationship between the corresponding gaze point and the display screen.

[0248] In this embodiment, the specific form of the eye information collected by the front-facing camera 310 is not limited. Optionally, the eye information collected by the front-facing camera 310 can be image information or video stream information. Depending on the processing performance of the first processing chip 320, the form of the eye information collected by the front-facing camera 310 can be different. For example, for a first processing chip 320 with high processing performance, the eye information collected by the front-facing camera 310 can be real-time video stream information, while for a first processing chip 320 with slightly weaker processing performance, the eye information collected by the front-facing camera 310 can be image information corresponding to multiple images.

[0249] like Figure 3 As shown in the embodiments of this application, the electronic device 300 also includes a software underlying framework 330 running on an operating system (OS), and a reading application 340 for displaying the text content of the e-book and the various illustrations therein. The specific type of the software underlying framework 330 is not limited. Optionally, the software underlying framework 330 can be the Swing framework.

[0250] In this embodiment, the underlying software framework 330 provides viewpoint tracking events and a second registration and callback interface to upper-layer applications. Correspondingly, the reading application 340 can request the underlying software framework 330 to register viewpoint tracking events for digital content and bind them to the second registration and callback interface. Based on this, during the user's e-book reading process, the first processing chip 320 can determine the relative position of the user's gaze point to the display screen through the viewpoint tracking events. Furthermore, the second registration and callback interface bound to the viewpoint tracking events can determine the screen coordinates of the user's gaze point on the display screen. Further, the second registration and callback interface can provide the corresponding screen coordinates to the reading application 340 so that the reading application 340 can determine whether the user is gazing at illustrations in the text content.

[0251] In this embodiment of the application, the reading application 340 integrates a typesetting engine, which is used to typeset the text content of the e-book and the position and layout of each illustration on the display screen to obtain the corresponding typesetting information.

[0252] like Figure 3 As shown, the reading application 340 also includes a first processing module 341, a second processing module 342, and a third processing module 343. The first processing module 341 is used to determine the relative positional relationship between the user's gaze point and each illustration in the text content based on the typesetting information obtained from the typesetting engine and the screen coordinates of the user's gaze point on the display screen provided by the software's underlying framework 330. The second processing module 342 is used to determine whether the user's gaze point falls within a first display area of ​​the first illustration on the display screen, based on the relative positional relationship between the user's gaze point and each illustration in the text content, thereby determining whether the user has a gaze intent towards the first illustration.

[0253] Optionally, if the second processing module 342 determines that the user intends to gaze at the first illustration, it can also display an interactive prompt message on the display screen indicating that the first illustration supports image enhancement processing, prompting the user to interact with the first illustration. Simultaneously, this triggers the third processing module 343 to perform image enhancement processing on the first illustration. Based on this, the third processing module 343 loads image enhancement data obtained from a cloud service according to the trigger command from the second processing module 342, and, in response to the user's interactive operation on the first illustration, displays the enhanced image of the first illustration on the display screen based on the loaded image enhancement data.

[0254] like Figure 3 As shown, the electronic device 300 also includes a preset interaction component 350 for responding to user touch operations. The preset interaction component 350 supports touch methods including, but not limited to, tap, long tap, swipe, rotation, and continuous taps. In this embodiment, the preset interaction component 350 pre-registers touch detection events to monitor whether the preset interaction component 350 is touched. Upon detecting that the preset interaction component 350 is touched, the system determines the user's touch method on the preset interaction component 350 and, through a first registration and callback interface bound to the registration detection event, selects a first illustration from the text content that requires image enhancement processing based on the corresponding touch method. Optionally, the first registration and callback interface can be provided by the underlying software framework 330 or by a separate software framework; this is not limited here.

[0255] In addition, to ensure that image enhancement processing can still be performed on various illustrations in the text content even when the electronic device 300 cannot communicate with the cloud server, such as... Figure 3As shown in the embodiment of this application, the electronic device 300 further includes an image enhancement module 360, wherein the image enhancement module 360 ​​integrates an artificial intelligence (AI) algorithm model, which is used to enhance the clarity of each illustration in the text content and generate a high-definition image corresponding to the illustration as an enhanced image.

[0256] Optionally, the specific method by which the image enhancement module 360 ​​performs image enhancement processing on each illustration in the text content can be found in the description of the following embodiments, and will not be detailed here.

[0257] It should be noted that, Figure 3 The device structure shown is for illustrative purposes only. In actual applications, the functional modules and components included in the electronic device 300 are not limited to this. Depending on the functional requirements, other functional modules or components may also be included, which will not be described in detail here.

[0258] Based on the above, the overall process of an electronic device acquiring digital content from a cloud server and displaying the first enhanced image is described below with reference to the accompanying drawings.

[0259] Figure 4a A signaling diagram illustrating the interaction between an electronic device and a cloud server is shown.

[0260] In this embodiment, after a user launches the reading application, they can view various ebooks provided by the cloud server. For ease of distinction, ebooks that the user has never read before are referred to as the first ebook. When the user selects the first ebook and opens it to read, such as... Figure 4a As shown in step (1), the electronic device can initiate a first retrieval request to the cloud server to request the digital content corresponding to the first e-book. Accordingly, as... Figure 4a As shown in steps (2) and (3), after receiving the first acquisition request, the cloud server can determine the digital content corresponding to the first e-book and send it to the electronic device.

[0261] Furthermore, after the electronic device receives the digital content returned by the cloud server, such as Figure 4a As shown in step (4), the electronic device can parse and process the corresponding digital content to obtain structured text data and text metadata corresponding to the text content of the e-book, as well as the original images and image metadata of each illustration in the text content of the e-book.

[0262] Furthermore, after the electronic device parses the above data, such as Figure 4aAs shown in step (5), the electronic device can use the typesetting engine to typeset the above structured text data and each original image according to the above data and screen display parameters, obtain the corresponding typesetting information, and display the corresponding text content and each illustration on the display screen for the user to read.

[0263] At the same time, such as Figure 4a As shown in steps (6)-(10), the reading application can register gaze tracking events for the corresponding digital content with the underlying software framework and bind a second registration and callback interface. During the user's use of the reading application, the front-facing camera can continuously collect the user's eye information. The first processing chip associated with the front-facing camera can trigger the gaze tracking event and identify the relative positional relationship between the user's gaze point and the display screen from the corresponding eye information. Then, the relative positional relationship between the user's gaze point and the display screen is input into the underlying software framework so that the screen coordinates of the user's gaze point on the display screen can be determined through the second registration and callback interface based on the aforementioned relative positional relationship.

[0264] Furthermore, such as Figure 4a As shown in steps (11)-(12), the reading application can obtain the screen coordinates returned by the second registration and callback interface, and determine the position information of each illustration in the e-book text content on the display screen based on the typesetting information generated by the typesetting engine. Based on this, the reading application can determine whether the user's gaze is located within the first display area occupied by the first illustration on the display screen, according to the screen coordinates of the user's gaze point on the display screen and the position information of each illustration in the e-book text content on the display screen. Furthermore, as... Figure 4a As shown in step (14), the reading application determines that the user has a gaze intent toward the first illustration when it determines that the user's gaze point is located within the first display area.

[0265] Optionally, considering that there may be some recognition error when the user's gaze point moves to the edge of the first display area, in order to improve recognition accuracy, a preset transition coefficient is pre-set in this embodiment to set a corresponding transition area for the display area occupied by each illustration on the display screen. Based on this, when the reading application determines whether the screen coordinates of the user's gaze point on the display screen are located within the first display area, it can determine whether the screen coordinates of the user's gaze point on the display screen are located within the first display area based on the screen coordinates of the user's gaze point on the display screen and the position information of the transition area corresponding to the first display area on the display screen. The specific form of the preset transition coefficient is not limited. Optionally, the preset transition coefficient can be a fixed pixel value or a proportional value. For example, the transition area corresponding to the display area of ​​each illustration on the display screen can be obtained by scaling the corresponding display area by a certain proportion.

[0266] In this embodiment, the method of setting a transition region for each illustration's display area on the screen is not limited. Optionally, for any illustration's display area, an inner transition region can be set based on a first preset transition coefficient on the inner side of the corresponding display area edge, and an outer transition region can be set based on a second preset transition coefficient on the outer side of the display area edge. Alternatively, a transition region can be set only on one side. The widths of the inner and outer transition regions corresponding to each display area can be the same or different, and correspondingly, the first and second preset transition coefficients can be the same or different.

[0267] For example, for larger illustrations, equally spaced inner and outer transition regions can be defined simultaneously on both the inner and outer sides of the corresponding display area edge. For smaller illustrations, a corresponding outer transition region can be defined only on the outer side of the corresponding display area edge, or a wider outer transition region can be defined on the outer side of the corresponding display area edge, while a narrower inner transition region is defined on the inner side of the corresponding display area edge. The specific configuration method can be determined according to the size of the display area, and will not be detailed here.

[0268] Based on this, such as Figure 4a As shown in step (13), before determining that the user's gaze point is located within the first display area, the reading application further sets a transition area for the first display area according to a preset transition coefficient. Accordingly, determining whether the screen coordinates of the user's gaze point on the display screen are located within the first display area means determining whether the user's gaze point is located within the first display area based on the screen coordinates of the user's gaze point on the display screen and the position information of the transition area corresponding to the first display area on the display screen.

[0269] In this embodiment, there is no limitation on the timing of setting the transition area for the first display area. Optionally, after the typesetting engine typeset the text content and illustrations of the e-book, the transition area can be set directly for the display area occupied by each illustration on the screen based on the typesetting information. In this way, based on the screen coordinates of the user's gaze point on the screen and the position information of the transition area corresponding to each display area on the screen, it can be determined whether the user's gaze point is located within the first display area. Alternatively, the transition area can be set for the first display area only after it has been determined that the user's gaze point is located within the first display area. Then, based on the screen coordinates of the user's gaze point on the screen and the position information of the transition area corresponding to the first display area on the screen, it can be determined whether the user's gaze point is located within the first display area.

[0270] Optionally, to improve recognition accuracy, in this embodiment, a corresponding preset sampling frequency is pre-set for the front-facing camera of the electronic device. Based on this, in the process of determining whether the screen coordinates of the user's gaze point on the display screen are located within the first display area, such as... Figure 4a As shown in steps (15)-(19), the front-facing camera is also used to continuously collect first eye information within a preset sampling frequency, and to determine multiple first screen coordinates of the user's gaze point on the display screen from the first eye information in accordance with steps (8)-(11). Furthermore, after the reading application obtains multiple first screen coordinates, it can determine whether all multiple first screen coordinates are located within the transition area corresponding to the first display area. Based on this, if all multiple first screen coordinates are located within the transition area corresponding to the first display area, such as... Figure 4a As shown in step (14), the reading application determines that the user's gaze point is located within the first display area.

[0271] Furthermore, to further determine whether the user intends to gaze at the first illustration, in this embodiment, a first preset duration is used to measure whether the user intends to gaze at the first illustration. Based on this, when the reading application determines that the screen coordinates of the user's gaze point on the display screen are within the first display area, it is also used to determine whether the user's first gaze duration on the first display area reaches the first preset duration, so that if the first gaze duration reaches the first preset duration, it is determined that the user intends to gaze at the first illustration.

[0272] In this embodiment, the interactive metadata generated by the cloud server for each original image is described using the example of metadata stored in image metadata. That is, when the electronic device obtains digital content from the cloud server, it also obtains the interactive metadata of each original image. Based on this, after parsing the digital content, the electronic device can identify the interactive metadata corresponding to each illustration from the image metadata. For ease of explanation, the interactive metadata corresponding to the first illustration is referred to as the first interactive metadata. The first interactive metadata indicates the interaction modes and enhancement methods supported by the first illustration. Depending on the interaction modes supported by the first illustration, the location where the user interacts with the first illustration differs; and depending on the enhancement methods supported by the first illustration, the display method of the first enhanced image corresponding to the first illustration differs.

[0273] Based on this, such as Figure 4a As shown in steps (20) and (21), when it is determined that the first gaze duration has reached the first preset duration, the reading application can prompt the user to perform interactive operations on different positions of the first illustration based on the interactive prompt information displayed on the display screen for the first interactive metadata.

[0274] At the same time, such as Figure 4aAs shown in steps (22)-(23), the reading application can obtain the image enhancement data corresponding to the first illustration from the cloud server, and load the corresponding image enhancement data after receiving the image enhancement data returned by the cloud server. Further, Figure 4a As shown in steps (25) and (26), when the reading application responds to the user's interactive operation on the first illustration, it can generate a first enhanced image corresponding to the first illustration based on the image enhancement data, and display the first enhanced image based on the first interactive metadata.

[0275] exist Figure 4a The interactive signaling diagram shown illustrates how a reading application retrieves image enhancement data corresponding to the first illustration from a cloud server when it determines that the user intends to gaze at the first illustration. However, in practical applications, this is not the only possibility. For example, considering the balance between storage space and loading speed, the reading application can also determine all illustrations related to chapters, sections, paragraphs, page numbers, etc., based on the user's reading progress, and retrieve the corresponding image enhancement data from the cloud server. This can be understood as the image enhancement data corresponding to the first illustration being pre-downloaded to the electronic device's local storage space. When the reading application responds to the user's interaction with the first illustration, it can directly load the corresponding image enhancement data from the electronic device's local storage space, resulting in higher efficiency.

[0276] Accordingly, Figure 4a The illustration uses the example of storing interactive metadata in image metadata, but in practical applications, it is not limited to this. For example, the interactive metadata corresponding to each original image can also be stored separately on a cloud server. In this case, electronic devices can retrieve the interactive metadata from the cloud server on demand.

[0277] It should be noted that, Figure 4a The steps shown are only used to illustrate one possible form of the interaction process between the electronic device and the cloud server, and are not limiting. In practical applications, the specific interaction process between the electronic device and the cloud server can be determined according to actual needs, provided that it conforms to the principles of the technical solution of this application, and will not be elaborated here.

[0278] based on Figure 4a The following is an exemplary description, with reference to the accompanying drawings, of how a reading application determines whether the user's gaze is within the first display area.

[0279] Figure 4b A schematic diagram showing the relationship between the screen coordinates of a user's gaze point on the display screen and the relative position of the first display area on the display screen is shown.

[0280] like Figure 4bAs shown, in this example, the solid-lined box represents the edge of the first display area. Using the top-left corner of the screen as the starting coordinate, assuming the reading application determines the top-left corner of the first display area on the screen as (X, Y) based on the layout information, and the width of the first display area corresponds to the pixel unit 'a', and the height corresponds to the pixel unit 'h', the reading application can determine the position information of the first display area on the display screen. Figure 4b In this context, an "×" indicates the user's gaze position on the display screen. Assume the second registration and callback interface determines the screen coordinates (x, y) corresponding to this gaze position. Based on this, the reading application can determine whether the user's gaze position is located within the first display area, using the position information of the first display area on the screen and the screen coordinates of the user's gaze. For example, Figure 4b Let's take an example where the user's gaze point is located within the first display area on the screen.

[0281] Further, optionally, to improve the accuracy of the reading application in determining whether the user's gaze is within the first display area, Figure 4c This diagram illustrates the relative positional relationship between the screen coordinates of the user's gaze point on the display screen and the transition area corresponding to the first display area on the display screen.

[0282] like Figure 4c As shown, in this example, the solid-lined box represents the edge of the first display area. The dashed-lined box outside the edge of the first display area and the edge of the first display area form the outer transition area corresponding to the first display area. The dashed-lined box inside the edge of the first display area and the edge of the first display area form the inner transition area corresponding to the first display area. Based on this, to reduce the recognition error during the movement of the user's gaze point at the edge of the first display area, in this embodiment, the sub-display area outside the inner transition area corresponding to the first display area is used as the first verification area to confirm that the user's gaze point is located within the first display area. Figure 4c The area marked with 'A' in the middle, encompassing the portion outside the outer transition area corresponding to the first display area, is designated as the second verification area to confirm that the user's gaze point is not within the first display area. Figure 4c The area marked with B in the text. That is to say, the user's gaze is determined to be within the first display area only if the screen coordinates of the user's gaze point on the display screen are within the first verification area.

[0283] Of course, the above examples are merely illustrative. Figure 4c The illustration shows an example of setting an inner transition area and an outer transition area for the first display area. In practical applications, this is not the only option; the specific setting method can be determined according to actual needs.

[0284] In this embodiment, the first interactive metadata includes interactive object information and display description information. The interactive object information includes the coordinate range corresponding to an independent interactive object in the first illustration, indicating whether the first illustration supports a global interactive mode or a local interactive mode. The display description information indicates whether the first illustration supports a global enhancement mode or a local enhancement mode. That is, if the coordinate range included in the interactive object information is the same as the coordinate range corresponding to the first illustration, then the first illustration supports a global interactive mode, meaning the first illustration is an independent interactive object. If the coordinate range included in the interactive object information is the same as the coordinate range corresponding to each sub-region or main element in the first illustration, then the first illustration supports a local interactive mode, meaning each sub-region or main element in the first illustration is displayed separately. Correspondingly, a global enhancement mode refers to the separate display of the enhanced image corresponding to the first illustration, while a local enhancement mode refers to the separate display of the enhanced image corresponding to each sub-region or main element in the first illustration.

[0285] It should be noted that the method for generating the first interactive metadata can be found in the description of the foregoing embodiments, and will not be repeated here.

[0286] The following description, in conjunction with the accompanying drawings, explains the process of determining the location where the user interacts with the first illustration based on the interaction mode supported by the first illustration, and the process of determining the display mode of the first enhanced image based on the enhancement mode supported by the first illustration.

[0287] For ease of explanation, the following embodiments are described from the perspective of an electronic device. For details regarding which functional module within the electronic device performs different processing procedures, please refer to the [reference needed]. Figure 3 and Figure 4a The relevant descriptions will not be distinguished in the following embodiments.

[0288] First, based on the interaction modes supported by the illustration, the locations where users interact with the first illustration are explained.

[0289] Example 1: The first interactive metadata indicates that the first illustration supports the global interactive mode.

[0290] In this example, the electronic device determines the coordinate range of the first illustration based on the interactive object information obtained from the first interactive metadata. Therefore, it determines that the first illustration supports a global interactive mode, meaning the first illustration is an independent interactive object. Based on this, when displaying interactive prompts, the corresponding interactive prompts can be directly displayed on the screen for the first illustration to prompt the user to interact with it. The display method of the interactive prompts is not limited. Optionally, the interactive prompts can be displayed inside the first display area or at the edge of the first display area. For example, the display state inside or at the edge of the first display area can be changed to serve as the interactive prompt for the first illustration. Correspondingly, the method of changing the display state inside or at the edge of the first display area is also not limited. Optionally, it may include, but is not limited to, color gradients, highlighting, glowing, flashing, etc. The specific method can be determined according to actual needs.

[0291] Based on this, when a user notices the interactive prompts displayed on the screen in relation to the first illustration, they can interact with any location within the first illustration. Furthermore, upon receiving a user interaction with the first illustration, the electronic device can generate a first enhanced image corresponding to the first illustration based on the acquired image enhancement data, and display the first enhanced image based on the enhancement method indicated by the first interactive metadata. The display method of the first enhanced image is not limited; optionally, it can be displayed directly within the first display area, or it can be displayed through a mask layer, area magnification, a separate window, etc. The specific display method can be determined according to actual needs.

[0292] Example 2: The first interactive metadata indicates that the first illustration supports a local interactive mode.

[0293] In this example, the electronic device determines the coordinate range of multiple sub-regions or main elements in the first illustration based on the interaction object information obtained from the first interaction metadata. Therefore, it determines that the first illustration supports a local interaction mode, meaning that the multiple sub-regions or main elements in the first illustration are independent interaction objects. Based on this, corresponding interaction prompts are displayed on the screen for each sub-region or main element to prompt the user to interact with it. For example, if the first illustration includes main elements such as clouds, seawater, land, and mountains, the electronic device determines that these main elements are independent interaction objects based on the interaction object information included in the first interaction metadata, and can display corresponding interaction prompts on the screen for each of these main elements.

[0294] In this embodiment, the display method of the interactive prompts is not limited. In one optional approach, the electronic device can simultaneously display the corresponding interactive prompts for each sub-region or main element on the display screen to prompt the user to interact with any sub-region or main element. In another optional approach, the electronic device can first determine the first sub-region or first main element that the user is looking at based on the screen coordinates of the user's gaze point on the display screen, and then display the corresponding interactive prompts for the first sub-region or first main element on the display screen. Here, the first sub-region can be any sub-region in the first illustration, and the first main element can be any main element in the first illustration. That is, if the user's gaze point moves between different sub-regions or different main elements, the interactive prompts corresponding to the gazed sub-region or main element can be displayed sequentially on the display screen.

[0295] Optionally, taking the first main element as an example, the electronic device can display corresponding interactive prompts within the area occupied by the first main element in the first illustration, or it can display corresponding interactive prompts at the edge of the first main element. The specific form of the prompts is not limited. For example, the display state of the area or edge occupied by the first main element in the first illustration can be changed as the corresponding interactive prompt. Accordingly, the changes in display state include, but are not limited to, color gradients, highlighting, glowing, flashing, etc., and the specific method can be determined according to actual needs. As another example, a preset identifier can be displayed within the area occupied by the first main element in the first illustration as the corresponding interactive prompt. The specific form of the preset identifier is not limited. Taking "cloud" as the first main element, the corresponding preset identifier can be text, graphics, symbols, etc., corresponding to "cloud," and the specific form can be determined according to actual needs.

[0296] Based on this, when a user notices the interactive prompts displayed for a sub-region or main element in the first illustration, they can interact with the first sub-region or the first main element. Correspondingly, upon receiving the user's interaction with the first sub-region or the first main element, the electronic device can generate a first enhanced image corresponding to the first illustration based on the acquired image enhancement data, and display the first enhanced image based on the enhancement method indicated by the first interactive metadata. The display method of the first enhanced image is not limited; optionally, it can be displayed directly within the first display area, or it can be displayed through a mask layer, area magnification, a separate window, etc. The specific display method can be determined according to actual needs.

[0297] Secondly, the display method of the first enhanced image is described according to the enhancement method supported by the first illustration.

[0298] In this example, for the case where the first illustration supports global enhancement, regardless of whether the illustration supports global interaction mode or local interaction mode, the user can perform interactive operations on any position of the first illustration. When the electronic device responds to the user's interactive operation on the first illustration, it generates an enhanced image of the first illustration based on the acquired image enhancement data, and displays the enhanced image on the display screen as the first enhanced image.

[0299] In this example, for the case where the first illustration supports local enhancement, the interaction mode supported by the first illustration is local interaction mode. The user can perform interactive operations on any sub-region or main element in the first illustration. When the electronic device responds to the user's interactive operation on the first sub-region or the first main element, it generates an enhanced image of the first sub-region or the first main element based on the acquired image enhancement data, and displays the first enhanced image on the display screen.

[0300] Based on this, the following description, in conjunction with the accompanying drawings, provides an exemplary account of how a user interacts with the first illustration in different ways, and how the first enhanced image is displayed based on the corresponding interactive operations.

[0301] Method 1: Users interact with the first illustration by gazing at it.

[0302] Based on the above, the front-facing camera of the electronic device can continuously collect the user's eye information. Based on the collected eye information, the electronic device can determine whether the user's gaze point is located within the first display area, thereby determining whether the user is continuously gazing at the first illustration.

[0303] In this embodiment, a second preset duration is used to measure whether the user interacts with the first illustration in a gaze manner. Based on this, when the electronic device determines that the user's first gaze duration on the first illustration has reached the first preset duration, it can also continuously acquire the user's second gaze duration on the first illustration, and when it determines that the second gaze duration has reached the second preset duration, it determines that the user has interacted with the first illustration.

[0304] The following section describes the display method of the first enhanced image in different interactive modes and enhancement methods.

[0305] 1) The first illustration supports global enhancement:

[0306] Figure 5aThe diagram illustrates a user's interactive operation on a first illustration, wherein the first illustration supports a global interaction mode, gaze 1 is the gaze corresponding to the user's first gaze duration on the first illustration reaching a first preset duration, and gaze 2 is the gaze corresponding to the user's second gaze duration on the first illustration reaching a second preset duration.

[0307] In this example, the electronic device determines that the first illustration supports a global interaction mode based on the interaction object information in the first interaction metadata. Based on this, the entire first illustration can be determined as an independent interaction object. Accordingly, the terminal prepares to determine that the first illustration supports a global enhancement mode based on the display description information in the first interaction metadata, and then generates an enhanced image of the entire first illustration as the first enhanced image based on the pre-acquired image enhancement data.

[0308] In this example, when the electronic device determines that the second gaze duration has reached a second preset duration, it can determine that the user has interacted with the first illustration, and based on this, displays the first enhanced image on the display screen. Figure 5a As shown, in the first enhanced image, the enhanced display content is all the display content in the first illustration.

[0309] Optionally, to prompt the user that the first illustration supports interactive operation, the electronic device may also display corresponding interactive prompt information for the first illustration after determining that the first gaze duration has reached a first preset duration, to prompt the user to continue gazing at the first illustration. For example... Figure 5a As shown, the electronic device can control the first illustration to emit edge light to prompt the user to continue looking at the first illustration. Of course, this prompting method is only an illustrative example and is not limited to this in actual applications.

[0310] Figure 5b Another schematic diagram of user interaction with the first sub-region is shown. The first illustration supports a local interaction mode. Line of sight 1 is the line of sight corresponding to the user's first gaze duration on the main element "cloud" reaching the first preset duration. Line of sight 2 is the line of sight corresponding to the user's second gaze duration on the main element "cloud" reaching the second preset duration.

[0311] In this embodiment, the electronic device determines that the first illustration supports a local interaction mode based on the interaction object information in the first interaction metadata. Based on this, the "cloud" as the main element can be determined as an independent interaction object. Accordingly, the terminal prepares to determine that the first illustration supports a global enhancement mode based on the display description information in the first interaction metadata. Then, based on the pre-acquired image enhancement data, it generates an enhanced image of the entire first illustration as the first enhanced image.

[0312] In this example, when the electronic device determines that the second gaze duration has reached a second preset duration, it can determine that the user has interacted with the main element of "clouds," and based on this, displays a first enhanced image on the screen. Figure 5b As shown, in the first enhanced image, the enhanced display content is all the display content in the first illustration.

[0313] Optionally, to prompt the user that the main element "cloud" supports interactive operation, the electronic device may also display corresponding interactive prompts for the main element "cloud" after determining that the first gaze duration has reached a first preset duration, to prompt the user to continue gazing at the main element "cloud". For example... Figure 5b As shown, the electronic device can control the main element of "cloud" to emit edge glow to prompt the user to continue looking at the main element. Of course, this prompting method is only an example and is not limited to this in actual applications.

[0314] 2) The first illustration supports local enhancement methods:

[0315] Figure 5c The diagram illustrates another user interaction with the first illustration, wherein the first illustration supports a local interaction mode, and line of sight 1 is the line of sight corresponding to the user's first gaze duration on the main element "cloud" reaching a first preset duration, and line of sight 2 is the line of sight corresponding to the user's second gaze duration on the main element "cloud" reaching a second preset duration.

[0316] In this embodiment, the electronic device determines that the first illustration supports a local interaction mode based on the interaction object information in the first interaction metadata. Based on this, the "cloud" as the main element can be determined as an independent interaction object. Accordingly, the terminal prepares to determine that the first illustration supports a local enhancement mode based on the display description information in the first interaction metadata. Then, based on the pre-acquired image enhancement data, it generates an enhanced image corresponding to the "cloud" as the main element, which serves as the first enhanced image.

[0317] In this example, when the electronic device determines that the second gaze duration has reached a second preset duration, it can determine that the user has interacted with the main element of "clouds," and based on this, a first enhanced image is displayed on the screen. Figure 5b As shown, in the first enhanced image, the enhanced display content is the main element corresponding to the "cloud" in the first illustration.

[0318] Optionally, to prompt the user that the main element "cloud" supports interactive operation, the electronic device may also display corresponding interactive prompts for the main element "cloud" after determining that the first gaze duration has reached a first preset duration, to prompt the user to continue gazing at the main element "cloud". For example... Figure 5c As shown, the electronic device can control the main element of "cloud" to emit edge glow to prompt the user to continue looking at the main element. Of course, this prompting method is only an example and is not limited to this in actual applications.

[0319] It should be noted that, in Figure 5a and Figure 5b In the image enhancement data loaded by the electronic device, the enhanced image data corresponding to the first inset is used. Figure 5c In this context, the image enhancement data loaded by the electronic device can be either the enhanced image data corresponding to the first illustration or the enhanced image data corresponding to the main element "cloud". The specific method is not limited.

[0320] Method 2: Users interact with the first illustration by clicking.

[0321] In this embodiment, when a user notices the interactive prompt information displayed on the screen, they can click on the first illustration to perform interactive operations on it.

[0322] The following section describes the display method of the first enhanced image in different interactive modes and enhancement methods.

[0323] 1) The first illustration supports global enhancement:

[0324] Figure 5d The diagram illustrates another user interaction with the first illustration, wherein the first illustration supports a global interaction mode, and gaze 3 is the gaze corresponding to the user's first gaze duration on the first illustration reaching a first preset duration.

[0325] In this example, the electronic device determines that the first illustration supports a global interaction mode based on the interaction object information in the first interaction metadata. Based on this, the entire first illustration can be determined as an independent interaction object. Accordingly, the electronic device determines that the first illustration supports a global enhancement mode based on the display description information in the first interaction metadata. Then, based on the pre-acquired image enhancement data, it generates an enhanced image of the entire first illustration as the first enhanced image.

[0326] In this example, upon receiving a user's click on the first illustration, the electronic device can determine that the user has interacted with the first illustration, and based on this, displays a first enhanced image on the screen. Figure 5d As shown, in the first enhanced image, the enhanced display content is all the display content in the first illustration.

[0327] Optionally, to prompt the user that the first illustration supports interactive operation, the electronic device may also display corresponding interactive prompt information for the first illustration after determining that the first gaze duration has reached a first preset duration, to prompt the user to click on the first illustration. For example... Figure 5d As shown, the electronic device can control the first illustration to prompt the user by illuminating its edges. Of course, this prompting method is only an illustrative example and is not limited to this in actual applications.

[0328] Figure 5e The illustration shows another user interaction with the first sub-region, where the first illustration supports a local interaction mode, and gaze 3 is the gaze corresponding to the user's first gaze duration on the main element "cloud" reaching a first preset duration.

[0329] In this embodiment, the electronic device determines that the first illustration supports a local interaction mode based on the interaction object information in the first interaction metadata. Based on this, the "cloud" as the main element can be determined as an independent interaction object. Accordingly, the terminal prepares to determine that the first illustration supports a global enhancement mode based on the display description information in the first interaction metadata. Then, based on the pre-acquired image enhancement data, it generates an enhanced image of the entire first illustration as the first enhanced image.

[0330] In this example, upon receiving a user's click on the "cloud" element, the electronic device determines that the user has interacted with the "cloud" element and, based on this, displays a first enhanced image on the screen. Figure 5e As shown, in the first enhanced image, the enhanced display content is all the display content in the first illustration.

[0331] Optionally, to prompt the user that the main element "cloud" supports interactive operation, the electronic device can also display corresponding interactive prompts for the main element "cloud" after determining that the first gaze duration has reached a first preset duration, prompting the user to click on the main element "cloud". For example... Figure 5b As shown, the electronic device can control the main element of "cloud" to prompt the user by making its edges glow. Of course, this prompting method is only an example and is not limited to this in actual applications.

[0332] 2) The first illustration supports local enhancement methods:

[0333] Figure 5f This illustrates another interactive intention for a user to interact with the first sub-region. The first illustration supports a local interaction mode, and gaze 3 is the gaze corresponding to the user's first gaze duration on the main element "cloud" reaching a first preset duration.

[0334] In this embodiment, the electronic device determines that the first illustration supports a local interaction mode based on the interaction object information in the first interaction metadata. Based on this, the "cloud" as the main element can be determined as an independent interaction object. Accordingly, the terminal prepares to determine that the first illustration supports a local enhancement mode based on the display description information in the first interaction metadata. Then, based on the pre-acquired image enhancement data, it generates an enhanced image corresponding to the "cloud" as the main element, which serves as the first enhanced image.

[0335] In this example, when the electronic device detects that the user has clicked on the main element "cloud," it determines that the user has interacted with the "cloud" element, and based on this, displays the first enhanced image on the screen. For example... Figure 5f As shown, in the first enhanced image, the enhanced display content is the main element corresponding to the "cloud" in the first illustration.

[0336] Optionally, to prompt the user that the main element "cloud" supports interactive operation, the electronic device can also display corresponding interactive prompts for the main element "cloud" after determining that the first gaze duration has reached a first preset duration, prompting the user to click on the main element "cloud". For example... Figure 5f As shown, the electronic device can control the main element of "cloud" to prompt the user by making its edges glow. Of course, this prompting method is only an example and is not limited to this in actual applications.

[0337] It should be noted that, in Figure 5d and Figure 5e In the image enhancement data loaded by the electronic device, the enhanced image data corresponding to the first inset is used. Figure 5f In this context, the image enhancement data loaded by the electronic device can be either the enhanced image data corresponding to the first illustration or the enhanced image data corresponding to the main element "cloud". The specific method is not limited.

[0338] Method 3: Users interact with the first illustration by touching preset interactive components.

[0339] In this embodiment, the electronic device pre-registers touch monitoring events for preset interactive components and binds them to a first registration and callback interface. When a user launches the reading application, the touch monitoring events are triggered simultaneously. Based on this, the user can use different touch methods to perform touch operations on the preset interactive components, instructing the first registration and callback interface to perform different controls on the various illustrations displayed on the screen, thereby enabling interactive operations on each illustration.

[0340] In this embodiment, the specific type of touch method is not limited. Optionally, it may include, but is not limited to, light touch, long touch, swipe, rotation, continuous light touch, etc. Based on this, when the touch monitoring event responds to the preset interactive component being touched, it can determine the user's touch method on the preset interactive component. Then, through the first registration and callback interface, the display state of the first illustration on the display screen is controlled according to the corresponding touch method, so as to realize different interactive operations on the first illustration.

[0341] For example, a user can tap a preset interactive component to turn a page in the e-book; a user can long-tap a preset interactive component to turn pages continuously; a user can perform touch operations such as sliding or rotating the preset interactive component to select different illustrations, sub-regions, or main elements on the display screen; a user can tap a preset interactive component repeatedly to select any illustration, sub-region, or main element on the display screen.

[0342] Of course, the above-described touch control methods and the control operations performed on the illustrations, sub-areas, or main elements displayed on the screen based on various touch control methods are merely illustrative examples and are not limited to these in practical applications. For example, if the electronic device supports a gyroscope, the illustrations, sub-areas, or main elements displayed on the screen can also be controlled by tapping the back panel of the electronic device. The specific control method can be set according to actual needs and will not be detailed here.

[0343] Based on this, when users notice the interactive prompts displayed on the screen, they can interact with the first illustration by touching the preset interactive components.

[0344] The following section describes the display method of the first enhanced image in different interactive modes and enhancement methods.

[0345] 1) The first illustration supports global enhancement:

[0346] Figure 5g An interactive diagram is shown, in which the first illustration supports a global interaction mode, and gaze 4 is the gaze corresponding to the user's first gaze duration on the first illustration reaching a first preset duration.

[0347] In this example, the electronic device determines that the first illustration supports a global interaction mode based on the interaction object information in the first interaction metadata. Based on this, the entire first illustration can be determined as an independent interaction object. Accordingly, the electronic device determines that the first illustration supports a global enhancement mode based on the display description information in the first interaction metadata. Then, based on the pre-acquired image enhancement data, it generates an enhanced image of the entire first illustration as the first enhanced image.

[0348] In this example, when the electronic device detects that a user has selected the first illustration by repeatedly tapping a preset interactive component, it can determine that the user has interacted with the first illustration, and based on this, displays a first enhanced image on the screen. Figure 5g As shown, in the first enhanced image, the enhanced display content is all the display content in the first illustration.

[0349] Optionally, to prompt the user that the first illustration supports interactive operation, the electronic device may also display corresponding interactive prompts for the first illustration after determining that the first gaze duration has reached a first preset duration, to prompt the user to select the first illustration. For example... Figure 5g As shown, the electronic device can control the first illustration to prompt the user by illuminating its edges. Of course, this prompting method is only an illustrative example and is not limited to this in actual applications.

[0350] Figure 5h Another schematic diagram of user interaction with the first sub-region is shown. The first illustration supports a local interaction mode. The line of sight 4 is the line of sight corresponding to the user's first gaze duration on the main element "cloud" reaching a first preset duration.

[0351] In this embodiment, the electronic device determines that the first illustration supports a local interaction mode based on the interaction object information in the first interaction metadata. Based on this, the "cloud" as the main element can be determined as an independent interaction object. Accordingly, the terminal prepares to determine that the first illustration supports a global enhancement mode based on the display description information in the first interaction metadata. Then, based on the pre-acquired image enhancement data, it generates an enhanced image of the entire first illustration as the first enhanced image.

[0352] In this example, when the electronic device detects that the user has selected the "cloud" main element by repeatedly tapping a preset interactive component, it determines that the user has interacted with the "cloud" main element, and based on this, displays a first enhanced image on the screen. Figure 5h As shown, in the first enhanced image, the enhanced display content is all the display content in the first illustration.

[0353] Optionally, to prompt the user that the main element "cloud" supports interactive operation, the electronic device may also display corresponding interactive prompts for the main element "cloud" after determining that the first gaze duration has reached a first preset duration, to prompt the user to select the main element "cloud". For example... Figure 5h As shown, the electronic device can control the main element of "cloud" to prompt the user by making its edges glow. Of course, this prompting method is only an example and is not limited to this in actual applications.

[0354] 2) The first illustration supports local enhancement methods:

[0355] Figure 5i Another schematic diagram of user interaction with the first sub-region is shown. The first illustration supports a local interaction mode. The line of sight 4 is the line of sight corresponding to the user's first gaze duration on the main element "cloud" reaching a first preset duration.

[0356] In this embodiment, the electronic device determines that the first illustration supports a local interaction mode based on the interaction object information in the first interaction metadata. Based on this, the "cloud" as the main element can be determined as an independent interaction object. Accordingly, the terminal prepares to determine that the first illustration supports a local enhancement mode based on the display description information in the first interaction metadata. Then, based on the pre-acquired image enhancement data, it generates an enhanced image corresponding to the "cloud" as the main element, which serves as the first enhanced image.

[0357] In this example, when the electronic device detects that the user has selected the "cloud" element by repeatedly tapping a preset interactive component, it can determine that the user has interacted with the "cloud" element, and based on this, a first enhanced image is displayed on the screen. For example... Figure 5i As shown, in the first enhanced image, the enhanced display content is the main element corresponding to the "cloud" in the first illustration.

[0358] Optionally, to prompt the user that the main element "cloud" supports interactive operation, the electronic device may also display corresponding interactive prompts for the main element "cloud" after determining that the first gaze duration has reached a first preset duration, to prompt the user to select the main element "cloud". For example... Figure 5i As shown, the electronic device can control the main element of "cloud" to prompt the user by making its edges glow. Of course, this prompting method is only an example and is not limited to this in actual applications.

[0359] It should be noted that, in Figure 5g and Figure 5h In the image enhancement data loaded by the electronic device, the enhanced image data corresponding to the first inset is used. Figure 5i In this context, the image enhancement data loaded by the electronic device can be either the enhanced image data corresponding to the first illustration or the enhanced image data corresponding to the main element "cloud". The specific method is not limited.

[0360] It should be noted that, in the above example, when the illustration supports a local interaction mode, each main element is treated as an independent interactive object. Similarly, when the illustration includes multiple sub-regions, the methods for users to interact with each sub-region and generate enhanced images, as well as the methods for displaying enhanced images on the screen when the illustration supports local enhancement, are similar to the implementation principle of treating each main element as an independent interactive object. For details, please refer to the explanation in the aforementioned example; they will not be elaborated upon here.

[0361] In the examples above, the user's direct interaction with the first illustration, the first sub-region, or the first main element is used as an example. In practical applications, since illustrations in e-book text content are usually static images, in order to distinguish between the two functions of "display" and "interaction," in an optional embodiment of this application, after obtaining the first illustration and the first interaction metadata, the electronic device can also determine the first illustration, the first sub-region, or the first main element as an independent interaction object based on the interaction object information in the first interaction metadata, and then generate a corresponding interactive interface to respond to the user's interaction. It can be understood that the user's interaction with the interactive interface is considered as an interaction with the first illustration, the first sub-region, or the first main element.

[0362] The following description, in conjunction with the accompanying drawings, provides an exemplary illustration of how the interactive interface is generated and displayed.

[0363] Based on the above, after acquiring and parsing digital content to obtain structured text data, text metadata, image metadata, and the original images used as illustrations, the electronic device typesets the structured text data and illustrations using a typesetting engine. The resulting text content and illustrations are then displayed on the screen. Specifically, when typeset each illustration, the typesetting engine first determines the display area occupied by each illustration on the screen, and then displays the corresponding illustration within that area.

[0364] Figure 5j A schematic diagram showing the first illustration on the display screen of an electronic device is shown.

[0365] In this example, when the typesetting engine typeset the first illustration, it determines the area occupied by the first illustration on the display screen as the first display area. Based on this, the typesetting engine displays the first illustration within the first display area on the display screen. Figure 5j As shown, the dashed area on the display screen is the first display area, and the grayscale image displayed therein is the first inset.

[0366] Assuming that in this example, the first illustration supports a global interaction mode and a global enhancement mode, based on this, the electronic device determines that the first illustration supports the global interaction mode according to the interaction object information in the first interaction metadata corresponding to the first illustration. Therefore, the electronic device can generate a corresponding first interactive interface for the entire first illustration to respond to user interaction operations. Further, upon receiving user interaction with the first interactive interface, the electronic device determines that the first illustration supports the global enhancement mode according to the display description information in the first interaction metadata corresponding to the first illustration. Then, based on pre-acquired image enhancement data, the electronic device generates a first enhanced image corresponding to the entire first illustration and displays the first enhanced image on the display screen.

[0367] It should be noted that the process of how the electronic device determines the interaction mode and enhancement method supported by the first illustration, and how to generate the first enhanced image, can be found in the relevant descriptions in the foregoing embodiments, and will not be repeated here.

[0368] The following describes how the first interactive interface is displayed.

[0369] In this embodiment, the implementation method of the first interactive interface is not limited; optionally, as shown below... Figure 5j As shown, the first interactive interface can be a mask layer covering the first illustration, serving as an interactive object for responding to user operations. The display method of the first interactive interface is not limited, and two optional methods are illustrated below.

[0370] In one alternative approach, the first interactive interface is invisible to the user. In this case, the first interactive interface serves only as an interactive object for responding to user interaction operations. Optionally, when the electronic device detects that the user's first gaze duration on the first illustration has reached a first preset duration, it can display interactive prompt information for the first illustration, as described in the above embodiments, instructing the first illustration to perform interactive operations on the first interactive interface. In this example, the timing of generating the first interactive interface is not limited. The first interactive interface can be generated directly after the layout engine typeset the first illustration and displays it in the first display area. Alternatively, the first interactive interface can be generated after the electronic device detects that the user's first gaze duration on the first illustration has reached the first preset duration. The specific generation timing can be determined according to actual needs.

[0371] In another alternative approach, the first interactive interface is visible to the user; for example, it can be a semi-transparent overlay. In this case, the first interactive interface can serve not only as an interactive object responding to user interactions but also as an optional form of interactive prompt information. This means that if the electronic device detects that the user's initial gaze duration on the first illustration has reached a first preset duration, the electronic device may not generate separate interactive prompt information for the first illustration. Therefore, in this example, when the layout engine typeset the first illustration and displays it within the first display area, the first interactive interface may not be generated for the first illustration. Instead, the first interactive interface is generated only after the electronic device detects that the user's initial gaze duration on the first illustration has reached the first preset duration, prompting the user to interact with the first illustration.

[0372] It should be noted that, for the sake of simplicity, the above example uses the first illustration supporting global interaction mode and global enhancement mode as an example. The implementation principle is similar for the case where the first illustration supports local interaction mode and global enhancement mode. That is, for the first illustration supporting global interaction mode, the electronic device can generate a corresponding interactive interface for the first illustration. After the user interacts with this interface, image enhancement processing of the first illustration can be triggered, and the enhanced image of the first illustration can be displayed. For the first illustration supporting local interaction mode, the electronic device can generate a corresponding interactive interface for the first sub-region or the first main element in the first illustration. After the user interacts with this interface, image enhancement processing of the first sub-region or the first main element can be triggered, and the enhanced image of the first sub-region or the first main element can be displayed. For the specific implementation process, please refer to the descriptions of the foregoing embodiments and examples, which will not be repeated here.

[0373] In the above embodiments, the process of image enhancement processing of the first illustration by the electronic device using image enhancement data generated by the cloud server is described in detail. It can be understood that, based on the powerful modeling capabilities of the cloud, the image enhancement data pre-generated by the cloud server for each illustration has a better enhancement effect, resulting in a higher clarity, richer visual format, and more realistic display effect in the enhanced image displayed on the electronic device. For example, in an illustration including adjacent land and sea, based on the image enhancement data generated by the cloud server for the corresponding illustration, the land portion is displayed with a higher-definition static effect in the final enhanced image displayed on the terminal, while the sea portion is displayed with a dynamic effect of flowing water ripples. As another example, in an illustration including an ocean with a small boat floating on its surface, based on the image enhancement data generated by the cloud server for the corresponding illustration, the sea portion is displayed with a dynamic effect of flowing water ripples in the final enhanced image displayed on the terminal, while the small boat is displayed with a dynamic effect of moving synchronously with the frequency of the ocean flow. For example, in an illustration, birds fly among mountains and clouds float on the mountainside. Based on the image enhancement data generated by the cloud server for the corresponding illustration, the final enhanced image displayed on the terminal shows birds flying among mountains and clouds floating on the mountainside, and so on.

[0374] In addition, refer to Figure 3 In this embodiment of the application, the electronic device can also enhance the clarity of the illustrations in the e-book based on its own internal image enhancement module. Even when the electronic device cannot obtain image enhancement data from the cloud service, it can still improve the display effect of each illustration in the text content of the e-book and enhance the user experience.

[0375] The process of the electronic device using its own image enhancement module to enhance the image of the first illustration will now be described with reference to the accompanying drawings.

[0376] Figure 6 An interactive signaling diagram is shown between various functional modules of an electronic device during the process of enhancing the sharpness of a first inset.

[0377] In this embodiment, the electronic device has obtained and parsed digital content from the cloud server, and the typesetting engine has typed the text content and illustrations of the e-book according to the parsing results, obtaining typesetting information. Based on this, during the user's reading of the e-book, the reading application can determine whether the screen coordinates of the user's gaze point on the display screen are located within the first display area corresponding to the first illustration on the display screen, based on the eye information continuously collected by the front-facing camera and the typesetting information obtained from the typesetting engine. The corresponding implementation process is as follows... Figure 6As shown in steps (1)-(14), the process of the electronic device obtaining digital content from the cloud server and parsing it, the typesetting engine typesetting the text content and illustrations of the e-book, and determining the screen coordinates of the user's gaze point on the display screen within the first display area can be referred to the description of the corresponding part in the foregoing embodiments, and will not be repeated here.

[0378] In this embodiment, the image enhancement module inside the electronic device integrates an AI algorithm model for enhancing the clarity of illustrations in an e-book and generating high-definition images corresponding to those illustrations. The AI ​​algorithm model relies on an NPU or GPU to perform operations such as parsing, recognition, noise reduction, and feature extraction on the illustrations. Optionally, the AI ​​algorithm model can work with the NPU or GPU provided by a first processing chip associated with the front-facing camera to enhance the clarity of the illustrations; alternatively, the image enhancement module can be independently associated with a second processing chip equipped with an NPU or GPU. Based on this, the AI ​​algorithm model can work with the NPU or GPU provided by the second processing chip to enhance the clarity of the illustrations. The specific implementation method can be determined according to the actual structure of the electronic device and is not limited here.

[0379] Optionally, regardless of whether the image enhancement module performs the sharpness enhancement processing operation through the NPU or GPU provided by the first processing chip or the second processing chip, for ease of description, the NPU or GPU provided by the first processing chip or the second processing chip is referred to as the image processing module.

[0380] Based on this, when the terminal application determines that the user's gaze point is located within the first display area on the screen, such as... Figure 6 As shown in steps (15) and (16), the user's first gaze duration on the first illustration can be continuously acquired, and when the first gaze duration is determined to reach the first preset duration, the interactive prompt information corresponding to the first illustration can be displayed on the screen to prompt the user to perform interactive operations on the first illustration.

[0381] Furthermore, such as Figure 6 As shown in steps (17) and (18), when the reading application responds to the user's interactive operation on the first illustration, it inputs the first illustration into the image processing module through an internal processing module (e.g., a third processing module) to trigger the image processing module and the image enhancement module to work together to perform image enhancement processing on the first illustration. Optionally, the user can interact with the first illustration by looking, clicking, or touching preset interactive components. For specific interaction methods, please refer to the description of the corresponding part in the foregoing embodiments, which will not be repeated here.

[0382] Based on this, such as Figure 6 As shown in step (19), after receiving the first illustration, the image processing module can parse and recognize the first illustration to identify the image parameters corresponding to the first illustration, such as, but not limited to, size, color, resolution, texture, etc. Further, as... Figure 6 As shown in steps (20)-(21), the image processing module inputs the identified image parameters and the first illustration into the image enhancement module. The image enhancement module can use the internally integrated AI algorithm model to perform main element recognition and sharpness enhancement processing on the first illustration according to the corresponding image parameters, and generate a second enhanced image corresponding to the first illustration.

[0383] Based on this, such as Figure 6 As shown in steps (22)-(24), the image enhancement module returns the generated second enhanced image to the reading application, and then displays the second enhanced image on the display screen to improve the display effect of the first illustration.

[0384] In this embodiment, the specific method of generating the second enhanced image by the image enhancement module is not limited. Depending on the capabilities of the AI ​​algorithm model, the method of generating the second enhanced image can also be different. For example, for a powerful AI algorithm model, the generated second enhanced image can be an enhanced image of animation or video, while for a less powerful AI algorithm model, the generated second enhanced image can be only a high-definition still image.

[0385] exist Figure 6 The example shown uses user interaction with the first illustration as an example. This can be understood as... Figure 6 In the example shown, the first illustration supports global interaction mode and global enhancement mode. The implementation principle is similar for the case where the first illustration supports local interaction mode and local enhancement mode.

[0386] For example, for the first illustration that supports local interactive modes, Figure 6 Steps (15)-(18) can be replaced by the following: when the reading application determines that the user's first gaze duration on the first sub-region or the first main element has reached a first preset duration, it displays interactive prompt information for the first sub-region or the first main element, and in response to the user's interactive operation on the first sub-region or the first main element, it inputs the first illustration to the image processing module and instructs for the second enhanced image corresponding to the first sub-region or the first main element. Based on this, Figure 6Steps (19)-(21) can be replaced by the image processing module identifying the image parameters corresponding to the first sub-region from the first illustration and inputting the corresponding image parameters into the image enhancement module, or identifying the image parameters corresponding to the first main element from the first illustration and inputting the corresponding image parameters into the image enhancement module; then, the image enhancement module performs sharpness enhancement processing on the first sub-region or the first main element based on the received image parameters to obtain the enhanced image corresponding to the first sub-region or the first main element, which serves as the second enhanced image.

[0387] It should be noted that, Figure 6 The steps shown are only for illustrating the principle of enhancing the clarity of the first illustration by the electronic device. In practical applications, the specific processing steps are not limited to these. For cases where the first illustration supports a local interactive mode, the methods for determining the supported interactive modes and enhancement methods, as well as the methods for user interaction with the first sub-region or the first main element and the display of the second enhanced image, can be found in the descriptions of the corresponding parts in the foregoing embodiments, and will not be repeated here.

[0388] Therefore, Figure 6 The clarity enhancement process shown can be used as a supplement to the aforementioned embodiments. For example, in the absence of network or network outage, the image enhancement service can continue to be provided to the user based on the image enhancement module locally supported by the electronic device, thus ensuring the user experience.

[0389] In summary, in this embodiment, during e-book reading, the front-facing camera of the electronic device continuously collects the user's eye information. A processing chip associated with the front-facing camera identifies the user's gaze point from this eye information, thereby determining the screen coordinates of the gaze point on the display screen. Simultaneously, the user can interact with any of the first illustrations in the e-book text at any time. Upon receiving this interaction, the electronic device displays a first enhanced image of the first illustration on the display screen based on image enhancement data corresponding to the first illustration obtained from a cloud server. The image enhancement data is generated by the cloud server after enhancing the elemental features of each main element in the first illustration. Therefore, the main elements included in the first enhanced image rendered based on the corresponding image enhancement data have a richer visual display style and a stronger sense of realism.

[0390] In summary, based on image enhancement data pre-generated on a cloud server, electronic devices can display the first enhanced image corresponding to the first illustration. This can transform obscure and abstract textual descriptions into rich and diverse enhanced images, which not only helps reduce the difficulty of understanding for users but also stimulates reading interest and enables immersive reading.

[0391] It is understood that the above embodiments are merely examples, and modifications can be made to the above embodiments in actual implementation. Those skilled in the art will understand that any modifications to the above embodiments that do not require creative effort fall within the protection scope of this application, and will not be described in detail in the embodiments.

[0392] Based on the same inventive concept, this application also provides a data processing method for enhancing display content. This method is used on a server, which will be referred to as the first server for clarity.

[0393] Figure 7 A flowchart of the data processing method for enhancing display content is shown, such as... Figure 7 The method includes:

[0394] S701. Obtain at least one original image from digital content, wherein the at least one original image is used as at least one illustration in text content, and the digital content is used to display the text content and at least one illustration on an electronic device.

[0395] S702. Recognize at least one original image to identify the element features of each subject element from each of the at least one original image.

[0396] S703. Generate image enhancement data for each original image based on the element features of each main element in each original image.

[0397] The implementation principle of the above-mentioned method steps by the first server will be explained below through specific embodiments.

[0398] Step S701: Obtain at least one original image from the digital content, the at least one original image being used as at least one illustration in the text content, the digital content being used to display the text content and at least one illustration on an electronic device.

[0399] Combination Figure 1a and Figure 1b In a digital reading scenario, the first server can obtain the raw data of ebooks from various ebook providers. After preprocessing the raw data, the corresponding digital content can be obtained. This digital content is used to display text content and at least one illustration on an electronic device. It includes structured text data, text metadata, at least one original image, and corresponding image metadata. Based on this, the first server can obtain at least one original image from the digital content to serve as at least one illustration in the text content. For detailed processing procedures involved in step S701, please refer to [link to relevant documentation]. Figure 1a and Figure 1b The relevant descriptions in the corresponding embodiments will not be repeated here.

[0400] Step S702: Recognize at least one original image to identify the element features of each subject element from each of the at least one original image.

[0401] In this embodiment, the first server can identify at least one original image using a multimodal large model to identify the element features of each subject element from each of the at least one original image. Optionally, for clarity, the multimodal large model used by the user in this method step is referred to as the first multimodal large model. Based on this, the first server can input a first prompt word and at least one original image into the first multimodal large model, instructing the first multimodal large model to identify at least one original image to identify the element features of each subject element from each of the at least one original image.

[0402] In this embodiment of the application, different interaction modes can be set for illustrations in the text content. Based on this, in order to generate corresponding image enhancement data for illustrations that support different interaction modes, the first server is also used to determine the interaction mode supported by the illustration before generating the image enhancement data corresponding to each original image. The interaction mode is either a global interaction mode or a local interaction mode. The global interaction mode is used to indicate that the illustration corresponding to each original image supports interactive operations, and the local interaction mode is used to indicate that multiple sub-regions or multiple main elements in the illustration corresponding to each original image support interactive operations.

[0403] Based on this, if the first server determines that the interaction mode supported by the illustration is a global interaction mode, in order to enable the electronic device to determine that each illustration is an independent interactive object, the first server is further configured to use the first coordinate range corresponding to each original image as interactive object information indicating the global interaction mode. Correspondingly, if the first server determines that the interaction mode supported by the illustration is a local interaction mode, in order to enable the electronic device to determine that multiple sub-regions or multiple main elements in each illustration are divided into independent interactive objects, the first server is further configured to perform an interactive object marking operation on each original image based on the element characteristics of each main element in each original image, so as to mark the coordinate ranges corresponding to the multiple sub-regions or multiple main elements in each original image respectively. Furthermore, based on the interactive object marking operation, a second coordinate range is determined for the multiple sub-regions or multiple main elements in each original image respectively, and the multiple second coordinate ranges corresponding to each original image are used as interactive object information indicating the local interaction mode.

[0404] Based on this, when the electronic device obtains the interactive object information corresponding to each original image, it can determine whether each illustration, multiple sub-regions within each illustration, or multiple main elements are independent interactive objects, allowing users to perform interactive operations. This results in diverse operation methods and a better user experience. Optionally, when the first server performs interactive object marking operations on each original image based on the element features of each main element in each original image, it can also be achieved through the first multimodal large model. Based on this, the first server can input a second prompt word into the first multimodal large model, instructing the first multimodal large model to perform interactive object marking operations on each original image based on the element features of each main element in each original image.

[0405] In this embodiment, corresponding enhancement methods can be set for illustrations in the text content, wherein the enhancement method is a global enhancement method or a local enhancement method. Based on this, in order for the electronic device to determine the enhancement method supported by the illustration, the first server can also determine the pre-set enhancement method for the illustrations in the text content, and generate display description information for indicating the corresponding enhancement method. Based on this, when the electronic device obtains the display description information, it can determine the display method of the enhanced image, and then display the enhanced image corresponding to each illustration according to the corresponding display method.

[0406] Optionally, the first server can also generate interactive metadata for each original image based on the display description information and the interactive object information corresponding to each original image. In this way, when the electronic device obtains the interactive metadata for each original image, it can retrieve the interactive object information and display description information based on the parsing results of the interactive metadata. Furthermore, based on the interactive object information, it determines the interaction mode supported by each illustration, to determine whether each illustration, multiple sub-regions within each illustration, or multiple main elements are independent interactive objects for responding to user interactions; and based on the display description information, it determines the display method of the enhanced image, and then displays the enhanced image corresponding to each illustration according to the corresponding display method. Further optionally, the first server can add the interactive metadata for each original image to the corresponding image metadata, or, based on the text metadata and image metadata, determine the association between each original image and structured text data, and add the interactive metadata for each original image to the corresponding text metadata based on the association. In this way, when the electronic device obtains the digital content corresponding to the e-book, it can simultaneously obtain the interactive metadata of each original image, which helps to save interaction times and facilitates processing.

[0407] Step S703: Generate image enhancement data for each original image based on the element features of each main element in each original image.

[0408] Optionally, when the first server generates image enhancement data corresponding to each original image based on the element features of each main element in each original image, it can also be achieved through a multimodal large model. For the sake of distinction, the multimodal large model used in this step is referred to as the second multimodal large model.

[0409] In this embodiment, the image enhancement data generated for each original image differs depending on the enhancement method supported by the illustration. Optionally, when the enhancement method supported by the illustration is a global enhancement method, the first server inputs the third prompt word, each original image, and the element features of each main element in each original image into the second multimodal large model, instructing the second multimodal large model to generate image enhancement data corresponding to each original image. Correspondingly, when the enhancement method supported by the illustration is a global enhancement method, the first server inputs the fourth prompt word, each original image, and the element features corresponding to each main element in each original image, as well as the marking results corresponding to the above marking operations into the second multimodal large model, instructing the second multimodal large model to generate image enhancement data corresponding to sub-regions or main elements in each original image.

[0410] It should be noted that, in the above embodiments, the specific forms and contents of the first and second multimodal large models used by the first server, as well as the first, second, third, and fourth prompt words involved therein, can be found in the foregoing embodiments regarding the use of the first, second, third, and fourth prompt words. Figure 2 The relevant explanations will not be repeated here.

[0411] In this embodiment of the application, in addition to executing the above-described method steps, the first server may optionally combine... Figure 1a and Figure 1b In the server architecture shown, with the first server as the master node server, the first server can also send digital content and image enhancement data to the second server, which acts as an edge server, so that the second server can communicate with the second electronic device and send the digital content and image enhancement data to the second electronic device.

[0412] Based on this, when the first server acts as an edge server, it can also communicate with electronic devices. Upon receiving a first acquisition request for digital content from the first electronic device, the first server can send the digital content to the first electronic device. Correspondingly, when the illustration supports a global enhancement method, upon receiving a second acquisition request from the electronic device for a first illustration in at least one illustration, the first server can determine the first original image as the first illustration, and then acquire the image enhancement data corresponding to the first original image and send it to the electronic device. When the illustration supports a global enhancement method, upon receiving a second acquisition request from the electronic device for a first illustration in at least one illustration, the first server can determine the first original image as the first illustration, and then acquire the image enhancement data corresponding to multiple sub-regions or multiple main elements in the first original image and send it to the electronic device.

[0413] In this embodiment, the first server, by executing the above method steps, can obtain image enhancement data that highly matches the element features of each main element in each original image. Thus, after the electronic device loads each image enhancement data, it can display the enhanced image corresponding to the content of the illustration. Based on this, by displaying the enhanced images corresponding to each illustration in the text content, the user can not only understand the text content more clearly and intuitively, reducing the difficulty of comprehension, but also enrich the form of digital reading and provide a better user experience.

[0414] Based on the same inventive concept, embodiments of this application also provide a method for processing display content in electronic devices.

[0415] Figure 8 A flowchart of a method for processing display content in an electronic device is shown, such as... Figure 8 As shown, the method includes:

[0416] S801. Obtain the first illustration and the first interactive metadata, wherein the first interactive metadata is used to instruct the generation of the first enhanced image corresponding to the first illustration;

[0417] S802, Display the first illustration on the display screen;

[0418] S803, Obtain the user's first gaze duration at the first illustration;

[0419] S804. If it is determined that the first gaze duration has reached the first preset duration, generate and display the first enhanced image according to the first interaction metadata.

[0420] The implementation process of the above-described method steps by an electronic device will be described below through specific embodiments.

[0421] For step S801, obtaining the first illustration and the first interactive metadata, the first interactive metadata is used to indicate the generation of the first enhanced image corresponding to the first illustration.

[0422] Combination Figure 1a and Figure 1b The interactive scenario diagram shown in this application embodiment indicates that the electronic device can obtain digital content from a cloud server. The digital content includes structured text data, text metadata, at least one original image, and image metadata. The at least one original image is used as at least one illustration in the text content. The at least one illustration includes a first illustration. The text metadata or image metadata includes first interactive metadata.

[0423] For step S802, the first illustration is displayed on the screen.

[0424] Based on the above, after acquiring the aforementioned digital content, the electronic device can use a typesetting engine to typeset the structured text data and at least one original image according to the text metadata and image metadata, obtaining corresponding typesetting information. Then, based on the typesetting information, it can display the text content corresponding to the structured text data and at least one illustration corresponding to the at least one original image on the display screen. Optionally, for the first illustration, the electronic device can determine the first display area occupied by the first illustration on the display screen based on the typesetting information, and then display the first illustration within the first display area.

[0425] In step S803, the user's first gaze duration at the first illustration is obtained.

[0426] Combination Figure 3 The diagram shows the structure of the electronic device. The electronic device can obtain the user's eye information through the front-facing camera. Based on this, the electronic device can determine the screen coordinates of the user's gaze point on the display screen from the user's eye information. Then, based on the screen coordinates and the first display area, it can obtain the user's first gaze duration on the first display area.

[0427] In this embodiment, the electronic device pre-registers gaze tracking events and binds a second registration and callback interface. Based on this, when the electronic device obtains the user's eye information through the front-facing camera, it can determine the relative position of the user's gaze point and the display screen based on the pre-registered gaze tracking events and the processing chip bound to the front-facing camera. Then, through the second registration and callback interface bound to the gaze tracking events, it determines the screen coordinates of the user's gaze point on the display screen according to the relative position relationship.

[0428] In this embodiment of the application, in order to improve the accuracy of determining the user's gaze intent based on the user's gaze point, see [link to relevant documentation]. Figure 4band Figure 4c After obtaining the typesetting information from the typesetting engine, the electronic device can determine the inner transition area corresponding to the first display area based on the size of the first illustration and the first preset transition coefficient. Further, based on the screen coordinates and the inner transition area, it determines whether the screen coordinates are within the first display area. If the screen coordinates are determined to be within the first display area, the user's first gaze duration on the first display area is obtained.

[0429] Optionally, to improve the accuracy of determining the screen coordinates of the user's gaze point on the display screen, a preset sampling frequency is set for the front-facing camera in this application example. Therefore, the electronic device can acquire the user's first eye information obtained by the front-facing camera within the preset sampling frequency. Correspondingly, the determination of the screen coordinates of the user's gaze point on the display screen is also based on multiple first screen coordinates determined from the first eye information acquired within the preset sampling frequency. Based on this, when the electronic device determines whether the screen coordinates are located within the first display area according to the screen coordinates and the inner transition area, it can first determine whether multiple first screen coordinates are all located within the sub-display area excluding the inner transition area within the first display area. If it is determined that multiple first screen coordinates are all located within the sub-display area, then it is determined that the screen coordinates are located within the first display area.

[0430] Accordingly, in addition to setting an inner transition region for the first display area, the electronic device can also set an outer transition region for the first display area. Based on this, after obtaining the typesetting information from the typesetting engine, the electronic device can determine the outer transition region corresponding to the first display area according to the size of the first illustration and the second preset transition coefficient. Accordingly, when determining whether the screen coordinates are within the first display area based on the screen coordinates and the inner transition region, if it is determined that multiple first screen coordinates are all outside the outer transition region, then it is determined that the screen coordinates are outside the first display area.

[0431] In this embodiment, by setting an inner transition area and an outer transition area for the first display area, when determining the screen coordinates of the user's gaze point on the display screen based on the acquired eye information, multi-frame smoothing can be achieved, which helps to reduce interference caused by the user moving their gaze at the edge of the first display area and improves calculation accuracy.

[0432] In step S804, if it is determined that the first gaze duration has reached the first preset duration, the first enhanced image is generated and displayed based on the first interaction metadata.

[0433] In this embodiment, the first interaction metadata includes interaction object information, wherein the interaction object information includes the coordinate range corresponding to an independent interaction object in the first illustration, to indicate the interaction mode supported by the first illustration, wherein the interaction mode supported by the first illustration is a global interaction mode or a local interaction mode. Based on this, when it is determined that the user's first gaze duration reaches a first preset duration, the electronic device can also determine the interaction mode supported by the first illustration based on the coordinate range included in the interaction object information in the first interaction metadata.

[0434] Optionally, if the coordinate range included in the interactive object information is the first coordinate range corresponding to the first illustration, then the first illustration is determined to support a global interactive mode. In this case, interactive prompts corresponding to the first illustration can also be displayed on the screen to prompt the user to interact with the first illustration. Correspondingly, if the coordinate range included in the interactive object information is the second coordinate range corresponding to at least one sub-region or at least one main element in the first illustration, then the first illustration is determined to support a local interactive mode. In this case, interactive prompts corresponding to at least one sub-region in the first illustration can be displayed on the screen to prompt the user to interact with the first sub-region, wherein the first sub-region is one of at least one sub-region; or, interactive prompts corresponding to at least one main element in the first illustration can be displayed on the screen to prompt the user to interact with the first main element, wherein the first main element is one of at least one main element.

[0435] Based on this, when the electronic device generates and displays the first enhanced image according to the first interaction metadata, it can obtain image enhancement data corresponding to the first illustration based on the user's interaction operation with the target interactive object. Then, based on the first interaction metadata and the image enhancement data, it generates and displays the first enhanced image. The target interactive object is not limited; it can be different depending on the interaction mode supported by the first illustration. For example, if the first illustration supports a global interaction mode, the target interactive object is the first illustration itself; if the first illustration supports a local interaction mode, the target interactive object is the first sub-region or the first main element within the first illustration.

[0436] Based on the description of the above embodiments, combined with Figures 5a-5i As can be seen, in this embodiment of the application, three methods are provided for interactive operation.

[0437] Based on this, when the electronic device acquires image enhancement data corresponding to the first illustration based on the user's interaction with the target interactive object, in one optional manner, it can continue to acquire the user's second gaze duration on the target interactive object, and if the second gaze duration reaches a second preset duration, it determines that the user has performed an interaction with the target interactive object and acquires the image enhancement data corresponding to the first illustration. In another optional manner, after seeing the interactive prompt information, the user directly clicks on the target interactive object. Based on this, the electronic device acquires the image enhancement data corresponding to the first illustration upon responding to the user's click on the target interactive object. In yet another optional manner, the user interacts with the target interactive object by touching a preset interactive component. Based on this, the electronic device acquires the image enhancement data corresponding to the first illustration upon responding to the user's interaction with the target interactive object by touching the preset interactive component.

[0438] In this embodiment, touch monitoring events are pre-registered for preset interactive components and bound to a first registration and callback interface. Based on this, the electronic device can respond to the user's touch operation on the preset interactive component based on the pre-registered touch monitoring events and determine the touch method corresponding to the touch operation. Then, through the first registration and callback interface bound to the touch monitoring events, the user's interaction operation on the target interactive object is determined according to the touch method. For the interaction methods that the user can use on the preset interactive component, and the operation methods corresponding to different interaction methods, please refer to the relevant descriptions in the foregoing embodiments, which will not be repeated here.

[0439] In this embodiment, the method of obtaining the image enhancement data corresponding to the first illustration is not limited. Optionally, when the electronic device responds to the user's interactive operation on the target interactive object, it can directly initiate a second acquisition request to the cloud server to obtain the image enhancement data corresponding to the first illustration from the cloud server. Alternatively, it can obtain the image enhancement data corresponding to the first illustration from a local preset storage space, wherein the image enhancement data in the preset storage space is obtained by the electronic device from the cloud server in advance. Regarding the timing of the electronic device obtaining the image enhancement data from the cloud server, please refer to the description of the corresponding part in the foregoing embodiments, which will not be repeated here.

[0440] In this embodiment, in addition to performing the above-described method steps, the electronic device can also send image enhancement data to the display device for the display device to display the first enhanced image, or directly display the first enhanced image based on the image enhancement data. Since the first illustration can support different enhancement methods, in order to determine how to display the first enhanced image, when generating and displaying the first enhanced image based on the first interactive metadata and the image enhancement data, the electronic device can also determine the first enhancement method supported by the first illustration based on the display description information included in the first interactive metadata, and then generate and display the first enhanced image based on the first enhancement method and the image enhancement data; wherein the first enhancement method is a global enhancement method or a local enhancement method.

[0441] Based on this, if the electronic device determines that the first enhancement method is a global enhancement method, it generates an enhanced image of the first illustration based on the image enhancement data, which is then displayed on the screen as the first enhanced image. Correspondingly, if the electronic device determines that the first enhancement method is a local enhancement method, it generates an enhanced image of the first sub-region or the first main element in the first illustration based on the image enhancement data, which is then displayed on the screen as the first enhanced image. Optionally, the first enhanced image is an image of the main element in the first illustration displayed in a micro-motion manner, resulting in a better display effect.

[0442] In other words, the first enhanced image may be an enhanced image of the entire first illustration, or it may be an enhanced image of a sub-region or main element in the first illustration. For users, the display method is more diversified. Especially for children's picture books, children can trigger various sub-regions or main elements in the first illustration to achieve the effect of "clicking and moving", which is more interesting and helps to stimulate children's reading interest.

[0443] Optionally, for details regarding the process of displaying the first enhanced image based on user interaction, please refer to the description of the corresponding section in the foregoing embodiments, and in conjunction with... Figures 5a-5i This allows for a more intuitive understanding of how the first enhanced image is enhanced in different interactive scenarios. Specific details will not be elaborated here.

[0444] Further optional, see Figure 3 As shown in the structure, the electronic device can also determine the first enhancement method supported by the first illustration from the first interactive metadata, perform sharpness enhancement processing on the first illustration based on the locally provided image enhancement processing module to obtain a second enhanced image, and then display the second enhanced image on the display screen.

[0445] In this embodiment, determining the display method of the second enhanced image is similar to determining the display method of the first enhanced image. If the electronic device determines that the first enhancement method is a global enhancement method, it performs sharpness enhancement processing on the first illustration based on the locally provided image enhancement processing module to obtain the corresponding enhanced image, which is used as the second enhanced image. Correspondingly, if the electronic device determines that the first enhancement method is a local enhancement method, it performs sharpness enhancement processing on the first sub-region or the first main element in the first illustration based on the locally provided image enhancement processing module to obtain the corresponding enhanced image, which is used as the second enhanced image.

[0446] In this embodiment, the electronic device utilizes its locally provided image enhancement processing function to generate a second enhanced image for the first illustration. This serves as a supplement to the aforementioned embodiments, allowing the electronic device to continue providing image enhancement processing services to the user based on its locally supported image enhancement module, even when data communication with the cloud server is impossible (e.g., no network, network outage), thus ensuring a good user experience. The specific functions and principles of the image enhancement module can be found in the foregoing embodiments. Figure 3 The corresponding explanations will not be repeated here.

[0447] In the embodiments of this application, the types of digital content include digital picture books, digital catalogs, e-books or digital games, etc., with a wider range of application scenarios. Different digital content can present different image enhancement effects during application, making it more interesting and displaying more diverse styles.

[0448] In this embodiment, the electronic device can determine whether the user intends to gaze at the first illustration based on the screen coordinates of the user's gaze point on the display screen. Furthermore, in response to the user's interactive operation on the first illustration, it can display a first enhanced image corresponding to the first illustration on the display screen based on the first interactive metadata and image enhancement data corresponding to the first illustration obtained from the cloud server. The image enhancement data is generated by the cloud server after enhancing the element features of each main element in the first illustration. Therefore, the main elements included in the first enhanced image rendered based on the corresponding image enhancement data have a richer visual display style and a stronger sense of realism. This not only transforms obscure and abstract textual descriptions into rich and diverse enhanced images, but also helps reduce the difficulty of understanding for users, stimulates reading interest, and achieves immersive reading.

[0449] The electronic device provided in the embodiments of this application is described below with reference to the accompanying drawings. The display content processing method provided in the embodiments of this application can be implemented in an electronic device.

[0450] For example, electronic devices can be mobile phones, tablets, wearable devices (such as smartwatches, smart bracelets, etc.), in-vehicle devices, laptops, personal computers (PCs), ultra-mobile personal computers (UMPCs), netbooks, smart home appliances (such as smart TVs), personal digital assistants (PDAs), distributed devices, in-vehicle systems, etc. This application does not limit the specific type of electronic device.

[0451] For example, electronic devices can be equipped with Portable electronic devices running Linux or other operating systems. Alternatively, the electronic device can also be other portable electronic devices, such as laptops. It should be understood that the electronic device can also be a desktop computer, etc.

[0452] Figure 9a This is a structural block diagram of an electronic device provided in an embodiment of this application.

[0453] like Figure 9a As shown, the electronic device 900 may include a processor 910, an external memory interface 920, an internal memory 921, a universal serial bus (USB) interface 930, a charging management module 940, a power management module 941, a battery 942, antenna 1, antenna 2, a mobile communication module 950, a wireless communication module 960, an audio module 970, a speaker 970A, a receiver 970B, a microphone 970C, a headphone jack 970D, a sensor module 980, buttons 990, a motor 991, an indicator 992, a camera 993, a display screen 994, and a subscriber identification module (SIM) card interface 995, etc. The sensor module 980 may include a pressure sensor 980A, a gyroscope sensor 980B, a barometric pressure sensor 980C, a magnetic sensor 980D, an accelerometer sensor 980E, a distance sensor 980F, a proximity sensor 980G, a fingerprint sensor 980H, a temperature sensor 980J, a touch sensor 980K, an ambient light sensor 980L, a bone conduction sensor 980M, etc.

[0454] Processor 910 may include one or more processing units, such as application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. In some embodiments, electronic device 900 may also include one or more processors 910.

[0455] The controller can be the nerve center and command center of the electronic device 900. The controller can generate operation control signals based on the instruction opcode and timing signals to control the fetching and execution of instructions.

[0456] The processor 910 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 910 is a cache memory. This memory can store instructions or data that the processor 910 has just used or that are used repeatedly. If the processor 910 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 910, and thus improves the efficiency of the electronic device 900.

[0457] The processor 910 can execute different operations by executing instructions to achieve different functions. These instructions may be, for example, instructions pre-stored in the memory before the electronic device leaves the factory, or instructions read from an app after the user installs a new app during use. This application embodiment does not limit the scope of these instructions.

[0458] In some embodiments, the processor 910 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0459] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 910 may include multiple I2C buses. The processor 910 can couple to the touch sensor 980K, charger, flash, camera 993, etc., through different I2C bus interfaces. For example, the processor 910 can couple to the touch sensor 980K through the I2C interface, enabling the processor 910 and the touch sensor 980K to communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 900.

[0460] The I2S interface can be used for audio communication. In some embodiments, the processor 910 may include multiple I2S buses. The processor 910 can be coupled to the audio module 970 via the I2S bus to enable communication between the processor 910 and the audio module 970. In some embodiments, the audio module 970 can transmit audio signals to the wireless communication module 960 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.

[0461] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 970 and the wireless communication module 960 can be coupled via the PCM bus interface. In some embodiments, the audio module 970 can also transmit audio signals to the wireless communication module 960 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0462] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 910 and the wireless communication module 960. For example, the processor 910 communicates with the Bluetooth module in the wireless communication module 960 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 970 can transmit audio signals to the wireless communication module 960 via the UART interface to enable music playback through Bluetooth headphones.

[0463] The MIPI interface can be used to connect the processor 910 to peripheral devices such as the display screen 994 and the camera 993. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 910 and the camera 993 communicate via the CSI interface to enable the electronic device 900 to capture images. The processor 910 and the display screen 994 communicate via the DSI interface to enable the electronic device 900 to display images.

[0464] The GPIO interface is configurable via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 910 to a camera 993, a display 994, a wireless communication module 960, an audio module 970, a sensor module 980, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0465] USB port 930 is a USB standard compliant interface, which can be a Mini USB port, Micro USB port, USB Type-C port, etc. USB port 930 can be used to connect a charger to charge electronic device 900, and can also be used for data transfer between electronic device 900 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.

[0466] It is understood that the interface connection relationships between the modules illustrated in this application are merely illustrative and do not constitute a structural limitation on the electronic device 900. In other embodiments, the electronic device 900 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0467] The charging management module 940 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 940 receives charging input from the wired charger via a USB interface 930. In some wireless charging embodiments, the charging management module 940 receives wireless charging input via the wireless charging coil of the electronic device 900. While charging the battery 942, the charging management module 940 can also supply power to the electronic device via the power management module 941.

[0468] The power management module 941 connects the battery 942, the charging management module 940, and the processor 910. The power management module 941 receives input from the battery 942 and / or the charging management module 940, providing power to the processor 910, internal memory 921, external memory, display 994, camera 993, and wireless communication module 960. The power management module 941 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 941 may also be located within the processor 910. In other embodiments, the power management module 941 and the charging management module 940 may be housed in the same device.

[0469] The wireless communication function of electronic device 900 can be implemented through antenna 1, antenna 2, mobile communication module 950, wireless communication module 960, modem processor, and baseband processor.

[0470] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 900 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.

[0471] The mobile communication module 950 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 900. The mobile communication module 950 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 950 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 950 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 950 may be housed in the processor 910. In some embodiments, at least some functional modules of the mobile communication module 950 and at least some modules of the processor 910 may be housed in the same device.

[0472] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to a speaker 970A, receiver 970B, etc.) or displays images or videos through a display screen 994. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 910 and may be housed in the same device as the mobile communication module 950 or other functional modules.

[0473] The wireless communication module 960 can provide solutions for wireless communication applications on the electronic device 900, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 960 can be one or more devices integrating at least one communication processing module. The wireless communication module 960 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 910. The wireless communication module 960 can also receive signals to be transmitted from processor 910, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0474] In some embodiments, antenna 1 of electronic device 900 is coupled to mobile communication module 950, and antenna 2 is coupled to wireless communication module 960, enabling electronic device 900 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).

[0475] In some embodiments, the wireless communication solution provided by the mobile communication module 950 enables the electronic device 900 to communicate with devices in the network (such as a cloud server), and the WLAN wireless communication solution provided by the wireless communication module 960 also enables the electronic device 900 to communicate with devices in the network (such as a cloud server). Thus, the electronic device 900 can transmit data with the cloud server.

[0476] Electronic device 900 implements display functions through a GPU, a display screen 994, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 994 and the application processor. The GPU performs mathematical and geometric calculations for graphics rendering. Processor 910 may include one or more GPUs, which execute program instructions to generate or modify display information. In this embodiment, the display screen 994 may include a monitor and a touch device. The monitor outputs display content to the user, and the touch device receives touch events input by the user on the display screen 994.

[0477] Display screen 994, also known as a screen, can be used to display images, videos, etc. Display screen 994 may include a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a mini LED, a micro LED, a micro-OLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 900 may include one or more display screens 994.

[0478] It should be understood that the display 994 may also include more components, such as a backlight panel and driving circuitry. The backlight panel provides a light source, and the display panel emits light based on this light. The driving circuitry controls whether the liquid crystal layer is transparent or opaque.

[0479] Electronic device 900 can achieve shooting function through ISP, camera 993, video codec, GPU, display 994 and application processor.

[0480] The ISP (Image Signal Processor) is used to process data fed back from the camera 993. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise and brightness. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 993.

[0481] Camera 993 is used to capture still images or videos. An object is projected onto a photosensitive element by an optical image generated through a lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, electronic device 900 may include at least two camera entities.

[0482] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 900 is selecting a frequency, the DSP is used to perform Fourier transforms on the frequency energy.

[0483] Video codecs are used to compress or decompress digital video. Electronic device 900 may support one or more video codecs. Thus, electronic device 900 can play or record video in various encoding formats, such as Moving Picture Experts Group (MPEG)-1, MPEG-2, MPEG-3, MPEG-4, etc.

[0484] An NPU (Neural Processing Unit) is a computational processor for neural networks (NNs). By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0485] The external memory interface 920 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 900. The external memory card communicates with the processor 910 through the external memory interface 920 to perform data storage functions. For example, music, photos, and videos can be stored on the external memory card.

[0486] The internal memory 921 can be used to store one or more computer programs, which include instructions. The processor 910 can execute the instructions stored in the internal memory 921, thereby causing the electronic device 900 to perform image processing methods (e.g., image cropping or image merging), various functional applications, and data processing as provided in some embodiments of this application. The internal memory 921 may include a program storage area and a data storage area. The program storage area may store the operating system; it may also store one or more application programs (e.g., a gallery, contacts, etc.). The data storage area may store data created during the use of the electronic device 900. Furthermore, the internal memory 921 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0487] Electronic device 900 can implement audio functions such as music playback and recording through audio module 970, speaker 970A, receiver 970B, microphone 970C, headphone jack 970D, and application processor.

[0488] The audio module 970 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 970 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 970 may be located in the processor 910, or some functional modules of the audio module 970 may be located in the processor 910.

[0489] The speaker 970A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. Electronic device 900 can listen to music or make hands-free calls through the speaker 970A.

[0490] The receiver 970B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 900 answers a telephone call or voice message, the receiver 970B can be brought close to the listener's ear to hear the voice.

[0491] Microphone 970C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 970C, inputting the sound signal into microphone 970C. Electronic device 900 may have at least one microphone 970C. In some embodiments, electronic device 900 may have two microphones 970C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 900 may have three, four, or more microphones 970C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0492] The 970D headphone jack is used to connect wired headphones. The 970D headphone jack can be a USB 930 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0493] Pressure sensor 980A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 980A can be disposed on display screen 994. There are many types of pressure sensors 980A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 980A, the capacitance between the electrodes changes, and electronic device 900 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 994, electronic device 900 detects the touch operation intensity based on pressure sensor 980A. Electronic device 900 can also calculate the touch position based on the detection signal from pressure sensor 980A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.

[0494] The gyroscope sensor 980B can be used to determine the motion attitude of the electronic device 900. In some embodiments, the gyroscope sensor 980B can determine the angular velocity of the electronic device 900 about three axes (i.e., the x, y, and z axes). The gyroscope sensor 980B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 980B detects the angle of the shake of the electronic device 900, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the electronic device 900 by moving in the opposite direction, thus achieving image stabilization. The gyroscope sensor 980B can also be used in navigation and motion-sensing game scenarios.

[0495] The barometric pressure sensor 980C is used to measure air pressure. In some embodiments, the electronic device 900 calculates altitude using the air pressure value measured by the barometric pressure sensor 980C to assist in positioning and navigation.

[0496] The magnetic sensor 980D includes a Hall sensor. The electronic device 900 can use the magnetic sensor 980D to detect the opening and closing of the flip cover. In some embodiments, when the electronic device 900 is a flip phone, the electronic device 900 can detect the opening and closing of the flip cover using the magnetic sensor 980D. Then, based on the detected opening and closing state of the cover or the flip cover, features such as automatic flip unlocking can be set.

[0497] The 980E accelerometer can detect the magnitude of acceleration of an electronic device 900 in various directions (typically three axes). When the electronic device 900 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of the electronic device, and is applicable to screen orientation switching, pedometers, and other applications.

[0498] A distance sensor 980F is used to measure distance. Electronic device 900 can measure distance via infrared or laser. In some embodiments, the electronic device 900 can utilize the distance sensor 980F for distance measurement to achieve fast focusing in shooting scenarios.

[0499] The proximity sensor 980G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The electronic device 900 emits infrared light outward through the LED. The electronic device 900 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that an object is near the electronic device 900. When insufficient reflected light is detected, the electronic device 900 can determine that no object is near the electronic device 900. The electronic device 900 can use the proximity sensor 980G to detect when a user holds the electronic device 900 close to their ear for a phone call, so as to automatically turn off the screen to save power. The proximity sensor 980G can also be used in holster mode and pocket mode for automatic unlocking and locking of the screen.

[0500] The ambient light sensor 980L is used to sense ambient light intensity. Electronic device 900 can adaptively adjust the brightness of display screen 994 based on the sensed ambient light intensity. The ambient light sensor 980L can also be used to automatically adjust white balance when taking photos. The ambient light sensor 980L can also work with proximity sensor 980G to detect whether electronic device 900 is in a pocket, preventing accidental touches.

[0501] The fingerprint sensor 980H is used to collect fingerprints. The electronic device 900 can utilize the characteristics of the collected fingerprints to achieve fingerprint unlocking, accessing application locks, taking photos with fingerprints, answering calls with fingerprints, etc.

[0502] Temperature sensor 980J is used to detect temperature. In some embodiments, electronic device 900 uses the temperature detected by temperature sensor 980J to execute a temperature handling strategy. For example, when the temperature reported by temperature sensor 980J exceeds a threshold, electronic device 900 performs thermal protection by reducing the performance of a processor located near temperature sensor 980J to reduce power consumption. In other embodiments, when the temperature is below another threshold, electronic device 900 heats battery 942 to prevent abnormal shutdown of electronic device 900 due to low temperature. In still other embodiments, when the temperature is below yet another threshold, electronic device 900 boosts the output voltage of battery 942 to prevent abnormal shutdown due to low temperature.

[0503] Touch sensor 980K, also known as a touch panel or touch-sensitive surface, can be located on display screen 994. The touch sensor 980K and display screen 994 together form a touchscreen, also called a "touchscreen." Touch sensor 980K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 994. In some embodiments, touch sensor 980K may also be located on the surface of electronic device 900, in a different position than display screen 994.

[0504] The bone conduction sensor 980M can acquire vibration signals. In some embodiments, the bone conduction sensor 980M can acquire vibration signals from the vibrating bone segments of the human vocal cords. The bone conduction sensor 980M can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor 980M can also be incorporated into headphones to form bone conduction headphones. The audio module 970 can parse the voice signals from the vibrating bone segments of the vocal cords acquired by the bone conduction sensor 980M to realize voice functionality. The application processor can parse heart rate information from the blood pressure signals acquired by the bone conduction sensor 980M to realize heart rate detection functionality.

[0505] Buttons 990 include a power button, volume buttons, etc. Buttons 990 can be mechanical buttons or touch-sensitive buttons. Electronic device 900 can receive button input and generate key signal inputs related to user settings and function control of electronic device 900.

[0506] Motor 991 can generate vibration alerts. Motor 991 can be used for incoming call vibration alerts and for touch vibration feedback. For example, different vibration feedback effects can be corresponding to touch operations applied to different applications (such as taking photos, playing audio, etc.). Motor 991 can also correspond to different vibration feedback effects for touch operations applied to different areas of the display screen 994. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.

[0507] Indicator 992 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.

[0508] The SIM card interface 995 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 995 to establish contact with the electronic device 900. The electronic device 900 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 995 can support Nano SIM cards, Micro SIM cards, and other SIM cards. Multiple cards can be inserted into the same SIM card interface 995 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 995 is also compatible with different types of SIM cards. The SIM card interface 995 is also compatible with external memory cards. The electronic device 900 interacts with the network through the SIM card to achieve functions such as voice calls and data communication. In some embodiments, the electronic device 900 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 900 and cannot be separated from it.

[0509] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 900. In other embodiments of this application, the electronic device 900 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0510] Figure 9b This is a software structure block diagram of an electronic device 900 provided in an embodiment of this application.

[0511] In some alternative implementations, the software system of the electronic device 900 can adopt the layered architecture of the Harmony system. The layered architecture of the Harmony system can include four layers, from bottom to top: kernel layer, system basic service layer, framework layer, and application layer.

[0512] The Harmony system employs a multi-kernel design, with options including the Linux kernel, the Harmony microkernel, and the Internet of Things operating system (LiteOS). This design allows devices with varying capabilities to choose the appropriate system kernel.

[0513] The kernel layer can include the kernel abstract layer (KAL), which provides basic kernel capabilities to other Harmony layers, such as process management, thread management, memory management, file system management, network management, and peripheral device management.

[0514] The system basic service layer is the core capability set of the Harmony system, which supports the Harmony system to provide application services through the framework layer in scenarios where multiple devices are deployed.

[0515] Optionally, the system basic service layer may include the following components: a set of basic system capability subsystems, a set of basic software service subsystems, a set of enhanced software service subsystems, the HarmonyOS driver framework (HDF) and hardware abstraction adaptation layer (HAL), a set of hardware service subsystems, and a proprietary hardware service subsystem.

[0516] The system's basic capability subsystem set provides fundamental capabilities for the operation, scheduling, and migration of distributed applications across multiple devices within the Harmony system. It comprises a distributed soft bus, distributed data management and file management, distributed task scheduling, the Ark runtime, and distributed security and privacy protection. The Ark runtime provides a multi-language runtime (C / C++ / JavaScript) and basic system class libraries, and also provides a runtime for Java programs statically generated using the Ark compiler (i.e., the parts of the application or framework layer developed using the Java language). The code for recording the call path in this embodiment can be implemented within the Ark runtime.

[0517] The basic software service subsystem provides common and general software services to the Harmony system. It comprises subsystems such as graphics and image processing, distributed media, distributed artificial intelligence (AI), multimodal input, mobile sensing development platform (MSDP) & device virtualization (DV), event notification, telephony services, and distributed digital image processing (DFX). Each subsystem can be customized to suit different device deployment environments, with each subsystem tailored to its functional granularity.

[0518] The Enhanced Software Services Subsystem suite provides the Harmony system with differentiated, enhanced software services for various devices. It comprises subsystems such as tablet business software, smart screen business software, in-vehicle business software, and Internet of Things (IoT) business software. The Enhanced Software Services Subsystem suite can be tailored to the deployment environment of different device types, at the subsystem level, and each subsystem can be further tailored at the functional level.

[0519] HDF and HAL form the foundation of Harmony's open hardware ecosystem, providing hardware capability abstraction at the top and development frameworks and runtime environments for various peripheral drivers at the bottom.

[0520] The hardware service subsystem suite provides common, adaptable hardware services for the Harmony system, consisting of subsystems for general sensors, location, power, USB, biometrics, and other hardware services. Each subsystem can be customized to suit different device deployment environments, with each subsystem allowing for functional tailoring.

[0521] The proprietary hardware service subsystem provides the Harmony system with differentiated hardware services for different devices, which may include proprietary hardware services for tablets, in-vehicle systems, wearables, and IoT devices. The proprietary hardware service subsystem can be tailored at the subsystem level, and each subsystem can be further tailored at the functional level.

[0522] The framework layer provides Harmony system applications with user program frameworks and meta-capability frameworks in multiple languages ​​such as Java / C / C++ / JavaScript, as well as multi-language framework APIs for various software and hardware services.

[0523] The application layer includes system applications and third-party applications, which can include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS. Applications in the Harmony system are built upon meta-services PA and FA.

[0524] For example, the technical solutions involved in the following embodiments can all be implemented in the electronic device 900 having the above-described hardware and software architecture.

[0525] Based on the same inventive concept, embodiments of this application also provide a data processing apparatus for enhancing display content. This data processing apparatus may include a processor and a memory. The memory may be used to store computer programs, and the processor may call and execute the computer programs stored in the memory to enable the data processing apparatus to implement the data method for enhancing display content described in the above embodiments.

[0526] It should be noted that since the principle by which this data processing device solves the problem is similar to the aforementioned data processing method for enhancing display content, the implementation of this data processing device can refer to the implementation of the aforementioned data processing method, and the repeated parts will not be described again.

[0527] Based on the same inventive concept, embodiments of this application also provide a display content processing apparatus, which may include a processor and a memory. The memory may be used to store computer programs, and the processor may call and execute the computer programs stored in the memory to enable the display content processing apparatus to implement the display content processing method described in the above embodiments.

[0528] It should be noted that since the principle by which this display content processing device solves the problem is similar to that of the aforementioned display content processing method, the implementation of this display content processing device can refer to the implementation of the aforementioned display content processing method, and the repeated parts will not be described again.

[0529] Based on the same inventive concept, this application also provides an electronic device. Figure 10 This is a structural block diagram of an electronic device provided in an embodiment of this application. Figure 10 As shown, the electronic device 110 may include: a processor 111, a memory 112, a display screen (not shown), etc. The memory 112 may be used to store programs / code pre-installed at the factory, or to store code for execution by the processor 111. The electronic device 110 may include components for executing the aforementioned... Figure 8 The components of the method shown are used to implement the steps in the corresponding method.

[0530] It should be understood that the electronic device 110 may also include other components, such as a transceiver. This electronic device 110 may correspond to the electronic device 900 provided in the above embodiments. The transceiver can be used to receive various operations input by the user. The processor 111 can be used to execute various processing procedures in the above method embodiments.

[0531] Based on the same inventive concept, embodiments of this application also provide a server, the structural framework of which is similar to... Figure 10 The structure shown is similar; please refer to [link / reference]. Figure 10 No duplicate illustrations will be shown.

[0532] In this embodiment, the server includes a processor and a memory. The memory can be used to store programs / code pre-installed at the time of the server's manufacture, or it can store code for the processor to execute. The server can include components for executing the above-described... Figure 7 The components of the method shown are used to implement the steps in the corresponding method.

[0533] It should be understood that the server may also include other components, such as transceivers. This server may correspond to the electronic device 900 provided in the above embodiments. The transceiver can be used to receive various operations input by the user. The processor can be used to execute various processing procedures in the above method embodiments.

[0534] Based on the same inventive concept, embodiments of this application also provide a system, which includes an electronic device and a server. The electronic device may include components for performing the above-described... Figure 8 The components of the method shown are used to implement the steps in the corresponding method. The server may include components for executing the above. Figure 7 The components of the method shown are used to implement the steps in the corresponding method. For details regarding the structure of the electronic device and server, and the implementation process of their corresponding functions, please refer to the descriptions of the corresponding parts in the foregoing embodiments; they will not be repeated here.

[0535] Based on the same inventive concept, embodiments of this application also provide a chip, which may include a processor for supporting devices to implement the methods described in the above embodiments of this application.

[0536] In some alternative implementations, the chip may also include memory for storing necessary programs and data for the device. The chip may consist of a single chip or may contain chips and other discrete components.

[0537] For example, the chip can be implemented using one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), application-specific integrated circuits (ASICs), system-on-chips (SoCs), central processors (CPUs), network processors (NPs), digital signal processors (DSPs), microcontrollers (MCUs), controllers, state machines, gate logic, discrete hardware components, any other suitable circuitry, or any combination of circuitry capable of performing the various functions described throughout this application.

[0538] The processor in the above embodiments can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above methods.

[0539] The memory in the above embodiments can be volatile memory or non-volatile memory, or it can include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0540] Based on the same inventive concept, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when run on an electronic device, causes the electronic device to implement the methods described in the above method embodiments.

[0541] Based on the same inventive concept, this application also provides a computer program product, which includes a computer program that, when run on an electronic device, causes the electronic device to implement the methods described in the above method embodiments.

[0542] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0543] If these functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0544] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, electronic devices, computer-readable storage media, computer program products, and chip systems described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0545] In the several embodiments provided in this application, it should be understood that the disclosed apparatus, electronic devices, and methods can be implemented in other ways. For example, the embodiments of the apparatus and electronic devices described above are merely illustrative; multiple components may be combined or integrated into another system, or some features may be omitted or not performed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection between devices or units through some interface, and may be electrical, mechanical, or other forms.

[0546] It should be understood that in the various embodiments of this application, the execution order of each step should be determined by its function and internal logic, and the size of each step number does not mean the order of execution, and does not constitute a limitation on the implementation process of the embodiments.

[0547] The various parts of this specification are described in a progressive manner. Similar or identical parts between the various embodiments can be referred to interchangeably. Each embodiment focuses on the differences from other embodiments. In particular, the embodiments of apparatus, electronic devices, chip systems, computer-readable storage media, and computer program products are basically similar to the method embodiments, so the descriptions are relatively simple; relevant details can be found in the descriptions within the method embodiments.

[0548] In this document, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Those skilled in the art will understand the specific meaning of the above terms in this application according to the specific circumstances. It should be noted that, without conflict, the embodiments and features in the embodiments of this application can be combined with each other. This application is not limited to any single aspect, nor to any single embodiment, nor to any combination and / or substitution of these aspects and / or embodiments. Moreover, each aspect and / or embodiment of this application can be used alone or in combination with one or more other aspects and / or embodiments thereof.

[0549] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them; although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and they should all be covered within the scope of this application.

Claims

1. A method for processing display content, characterized in that, Applied to electronic devices, the method includes: Obtain a first illustration and first interactive metadata, wherein the first interactive metadata is used to instruct the generation of a first enhanced image corresponding to the first illustration; Display the first illustration on the screen; Obtain the user's first gaze duration at the first illustration; If it is determined that the first gaze duration has reached the first preset duration, the first enhanced image is generated and displayed based on the first interaction metadata.

2. The method according to claim 1, characterized in that, The first interactive metadata includes interactive object information, which includes the coordinate range of an independent interactive object in the first illustration. The interactive object information is used to indicate the interactive modes supported by the first illustration.

3. The method according to claim 2, characterized in that, If it is determined that the first gaze duration has reached a first preset duration, the method further includes: If the coordinate range is the first coordinate range corresponding to the first illustration, it is determined that the first illustration supports the global interaction mode; In the global interaction mode, interactive prompts corresponding to the first illustration are displayed on the screen to prompt the user to interact with the first illustration.

4. The method according to claim 2, characterized in that, If it is determined that the first gaze duration has reached a first preset duration, the method further includes: If the coordinate range is the second coordinate range corresponding to at least one sub-region or at least one main element in the first illustration, then the first illustration is determined to support a local interactive mode. In the local interaction mode, interactive prompts corresponding to at least one sub-region in the first illustration are displayed on the screen to prompt the user to perform an interactive operation on the first sub-region, where the first sub-region is one of the at least one sub-region; or... In the local interaction mode, interactive prompts corresponding to at least one main element in the first illustration are displayed on the display screen to prompt the user to perform an interactive operation on the first main element, which is one of the at least one main element.

5. The method according to claim 3 or 4, characterized in that, Based on the first interactive metadata, generate and display the first enhanced image, including: Based on the user's interactive operation on the target interactive object, image enhancement data corresponding to the first illustration is obtained, wherein the target interactive object is one of the first illustration, the first sub-region in the first illustration, or the first main element; The first enhanced image is generated and displayed based on the first interactive metadata and the image enhancement data.

6. The method according to claim 5, characterized in that, Based on the user's interaction with the target interactive object, image enhancement data corresponding to the first illustration is obtained, including: Obtain the second gaze duration of the user on the target interactive object; If the second gaze duration is determined to reach the second preset duration, it is determined that the user has performed an interactive operation on the target interactive object, and the image enhancement data corresponding to the first illustration is obtained.

7. The method according to claim 5, characterized in that, Based on the user's interaction with the target interactive object, image enhancement data corresponding to the first illustration is obtained, including: In response to the user's click operation on the target interactive object, image enhancement data corresponding to the first illustration is obtained.

8. The method according to claim 5, characterized in that, Based on the user's interaction with the target interactive object, image enhancement data corresponding to the first illustration is obtained, including: In response to the user's interactive operation on the target interactive object through a preset touch interactive component, image enhancement data corresponding to the first illustration is obtained.

9. The method according to claim 8, characterized in that, Responding to the user's interactive operation on the target interactive object through touch of a preset interactive component includes: Based on pre-registered touch monitoring events, respond to the user's touch operation on the preset interactive component and determine the touch method corresponding to the touch operation; The user's interaction with the target interactive object is determined through the first registration and callback interface bound to the touch monitoring event, based on the touch method.

10. The method according to any one of claims 5-9, characterized in that, The first interactive metadata includes display description information, which is used to indicate the enhancement methods supported by the first illustration; Generate and display the first enhanced image based on the first interactive metadata and the image enhancement data, including: Based on the display description information, determine the first enhancement method supported by the first illustration; The first enhanced image is generated and displayed based on the first enhancement method and the image enhancement data.

11. The method according to claim 10, characterized in that, Generate and display the first enhanced image based on the first enhancement method and the image enhancement data, including: If the first enhancement method is a global enhancement method, an enhanced image of the first illustration is generated based on the image enhancement data, and used as the first enhanced image; The first enhanced image is displayed on the display screen.

12. The method according to claim 10, characterized in that, Generate and display the first enhanced image based on the first enhancement method and the image enhancement data, including: If the first enhancement method is a local enhancement method, an enhanced image of the first sub-region or the first main element in the first illustration is generated based on the image enhancement data, and used as the first enhanced image; The first enhanced image is displayed on the display screen.

13. The method according to any one of claims 5-12, characterized in that, Obtaining the image enhancement data corresponding to the first illustration includes: Obtain the image enhancement data corresponding to the first illustration from the cloud server; Alternatively, image enhancement data corresponding to the first illustration can be obtained from a local preset storage space, and the image enhancement data is obtained in advance from the cloud server.

14. The method according to any one of claims 10-13, characterized in that, The method further includes: According to the first enhancement method, the first illustration is subjected to sharpness enhancement processing to obtain a second enhanced image; The second enhanced image is displayed on the display screen.

15. The method according to claim 14, characterized in that, According to the first enhancement method, the first illustration is subjected to sharpness enhancement processing to obtain a second enhanced image, including: If the first enhancement method is a global enhancement method, the first illustration is subjected to sharpness enhancement processing to obtain the corresponding enhanced image, which is used as the second enhanced image.

16. The method according to claim 14, characterized in that, According to the first enhancement method, the first illustration is subjected to sharpness enhancement processing to obtain a second enhanced image, including: If the first enhancement method is a local enhancement method, the first sub-region or the first main element in the first illustration is subjected to sharpness enhancement processing to obtain the corresponding enhanced image, which is used as the second enhanced image.

17. The method according to any one of claims 1-16, characterized in that, Obtain the metadata of the first illustration and the first interaction, including: Digital content is obtained from a cloud server. The digital content includes structured text data, text metadata, at least one original image, and image metadata. The at least one original image is used as at least one illustration in the text content. The at least one illustration includes the first illustration. The text metadata or the image metadata includes the first interactive metadata. Based on the typesetting engine, the text structure data and the at least one original image are typed according to the text metadata and the image metadata to obtain the corresponding typesetting information; Based on the layout information, the text content and the at least one illustration are displayed on the screen.

18. The method according to claim 17, characterized in that, Displaying the first illustration on the screen includes: Based on the layout information, the first display area occupied by the first illustration on the display screen is determined; The first illustration is displayed within the first display area.

19. The method according to claim 18, characterized in that, Obtaining the user's first gaze duration at the first illustration includes: Determine the screen coordinates of the user's gaze point on the display screen; Based on the screen coordinates and the first display area, the user's first gaze duration on the first illustration is obtained.

20. The method according to claim 19, characterized in that, Determining the screen coordinates of the user's gaze point on the display screen includes: Obtain the user's eye information; Based on pre-registered gaze tracking events, the relative positional relationship between the user's gaze point and the display screen is determined from the eye information; The screen coordinates of the user's gaze point on the display screen are determined based on the relative position relationship through the second registration and callback interface bound to the gaze point tracking event.

21. The method according to claim 20, characterized in that, Based on the screen coordinates and the first display area, the first gaze duration of the user on the first illustration is obtained, including: Based on the size of the first illustration and the first preset transition coefficient, the inner transition area corresponding to the first display area is determined; Based on the screen coordinates and the inner transition area, determine whether the screen coordinates are located within the first display area; If the screen coordinates are determined to be within the first display area, the user's first gaze duration at the first illustration is obtained.

22. The method according to claim 21, characterized in that, The eye information includes first eye information obtained from the user within a preset sampling frequency, and the screen coordinates include multiple first screen coordinates determined based on the first eye information. Determining whether the screen coordinates are located within the first display area based on the screen coordinates and the inner transition area includes: Determine whether all of the plurality of first screen coordinates are located within a sub-display area excluding the inner transition area within the first display area; If it is determined that all of the plurality of first screen coordinates are located within the sub-display area, then the screen coordinates are determined to be located within the first display area.

23. The method according to claim 22, characterized in that, The method further includes: Based on the size of the first illustration and the second preset transition coefficient, the outer transition area corresponding to the first display area is determined; If it is determined that all of the plurality of first screen coordinates are located outside the outer transition region, then it is determined that the screen coordinates are located outside the first display area.

24. The method according to any one of claims 1-23, characterized in that, The first enhanced image is an image in which the main elements in the first illustration are displayed in a micro-motion manner.

25. The method according to any one of claims 17-24, characterized in that, The types of digital content include digital picture books, digital catalogs, e-books, or digital games.

26. A data processing method for enhancing displayed content, characterized in that, Applied to a server, the method includes: Acquire at least one original image from digital content, said at least one original image being used as at least one illustration in text content, said digital content being used to display said text content and said at least one illustration on an electronic device; The at least one original image is identified to identify the element features of each main element from each of the at least one original image; Based on the element features of each main element in each original image, image enhancement data corresponding to each original image is generated.

27. The method according to claim 26, characterized in that, The at least one original image is identified to determine the element features of each subject element from each of the at least one original image, including: The first prompt word and the at least one original image are input into the first multimodal large model, which is then instructed to recognize the at least one original image to identify the element features of each subject element from each of the at least one original image.

28. The method according to claim 27, characterized in that, The method further includes: The illustration is determined to support a global interaction mode, which is used to indicate that the illustration corresponding to each original image supports interactive operations. The first coordinate range corresponding to each original image is used as the interaction object information indicating the global interaction mode.

29. The method according to claim 27, characterized in that, The method further includes: The illustration is determined to support a local interaction mode, which is used to indicate that multiple sub-regions or multiple main elements in the illustration corresponding to each original image support interactive operations. Based on the element features of each main element in each original image, an interactive object marking operation is performed on each original image to mark the coordinate ranges corresponding to multiple sub-regions or multiple main elements in each original image.

30. The method according to claim 29, characterized in that, Based on the element features of each main element in each original image, an interactive object marking operation is performed on each original image, including: The second prompt word is input into the first multimodal large model, instructing the first multimodal large model to perform interactive object marking operations on each original image based on the element features of each main element in each original image.

31. The method according to claim 29 or 30, characterized in that, The method further includes: Based on the interactive object marking operation, determine the second coordinate range for marking multiple sub-regions or multiple main elements in each original image; The multiple second coordinate ranges corresponding to each original image are used as interactive object information to indicate the local interactive mode.

32. The method according to claim 28 or 31, characterized in that, The method further includes: Determine the enhancement methods supported by the illustrations; Generate display description information to indicate the enhancement method; Based on the display description information and the interactive object information corresponding to each original image, generate interactive metadata corresponding to each original image.

33. The method according to any one of claims 28-31, characterized in that, Based on the element features of each main element in each original image, image enhancement data corresponding to each original image is generated, including: The illustration is determined to support a global enhancement method. The third prompt word, each original image, and the element features of each main element in each original image are input into the second multimodal large model. The second multimodal large model is instructed to generate image enhancement data corresponding to each original image. The global enhancement method is used to instruct the enhanced image of each illustration to be displayed separately.

34. The method according to claim 33, characterized in that, The method further includes: Upon receiving a second acquisition request from the electronic device for a first illustration among the at least one illustrations, a first original image is determined as the first illustration; The image enhancement data corresponding to the first original image is obtained and sent to the electronic device.

35. The method according to any one of claims 29-31, characterized in that, Based on the element features of each main element in each original image, image enhancement data corresponding to each original image is generated, including: The illustration is determined to support local enhancement. The fourth prompt word, each original image, the element features of each main element in each original image, and the marking result corresponding to the interactive object marking operation are input into the second multimodal large model. The second multimodal large model is instructed to generate image enhancement data corresponding to the sub-regions or main elements in each original image. The local enhancement method is used to instruct the enhanced images of the sub-regions or main elements in each illustration to be displayed separately.

36. The method according to claim 35, characterized in that, The method further includes: Upon receiving a second acquisition request from the electronic device for a first illustration among the at least one illustrations, a first original image is determined as the first illustration; Image enhancement data corresponding to multiple sub-regions or multiple main elements in the first original image are obtained and sent to the electronic device.

37. The method according to any one of claims 32-36, characterized in that, The digital content includes structured text data, text metadata, the at least one original image, and corresponding image metadata; the method further includes: Add the interactive metadata corresponding to each original image to the corresponding image metadata; or... Based on the text metadata and the image metadata, determine the association between each original image and the structured text data; Based on the association, the interactive metadata corresponding to each original image is added to the corresponding text metadata.

38. The method according to claim 37, characterized in that, The method further includes: Upon receiving a first request from the electronic device for the digital content, the digital content is sent to the electronic device.

39. An electronic device, characterized in that, include: One or more processors; One or more memory units; And one or more computer programs, wherein the one or more computer programs are stored in the one or more memories, the one or more computer programs including instructions that, when executed by the one or more processors, cause the electronic device to perform the method as described in any one of claims 1 to 25.

40. A server, characterized in that, include: One or more processors; One or more memory units; And one or more computer programs, wherein the one or more computer programs are stored in the one or more memories, the one or more computer programs including instructions that, when executed by the one or more processors, cause the server to perform the method as described in any one of claims 26 to 38.

41. A system, characterized in that, The system includes: An electronic device for performing the method as described in any one of claims 1 to 25; A server for performing the method as described in any one of claims 26 to 38.

42. A chip, characterized in that, The chip stores instructions that, when executed, implement the method as described in any one of claims 1 to 25, or the method as described in any one of claims 26 to 38.

43. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1-38.

44. A computer program product, characterized in that, The computer program product includes a computer program that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1-38.