Image processing method and device and electronic equipment

By using TMS slicing algorithm and geographic information system technology to slice ancient calligraphy and paintings, and combining the OpenLayers engine and custom projection, the problems of high cost and insufficient interactivity in the digitization of ancient calligraphy and paintings are solved, and smooth loading and enhanced interactivity are achieved on low-end devices.

CN121883271APending Publication Date: 2026-04-17AERIAL PHOTOGRAMMETRY & REMOTE SENSING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AERIAL PHOTOGRAMMETRY & REMOTE SENSING CO LTD
Filing Date
2025-12-31
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies have shortcomings in terms of display and interaction. In particular, the digitization of ancient calligraphy and paintings is costly, requires high equipment performance, and lacks interactivity and dissemination capabilities, making it difficult to achieve high-fidelity restoration and widespread application on mobile devices.

Method used

The target digital image is sliced ​​using the TMS slicing algorithm. Combined with the tile technology of the geographic information system and custom projection, multi-level sliced ​​images are generated and rendered using the OpenLayers engine. It supports voice generation and hotspot interaction, achieving smooth image display and interaction.

Benefits of technology

It significantly improves loading speed and responsiveness, reduces device performance requirements, supports smooth loading on low-end devices, enhances interactivity and dissemination, and provides a good user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883271A_ABST
    Figure CN121883271A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and device and electronic equipment, and the method comprises the steps: carrying out the slicing of a target digital image, and obtaining a plurality of first slice images; based on a TMS slicing algorithm, determining the number of levels corresponding to the target digital image and a second slice image corresponding to each level; and determining and loading the target slice image according to the image operation instruction to obtain a rendered image. According to the mode, only part of target slice images needed by a user are loaded each time through a slice technology, the whole target digital image is prevented from being downloaded, the loading speed and the response smoothness are remarkably improved, original image resolution details are reserved to the maximum extent through high-sampling multi-level slices, and good experience can be guaranteed even in the environment of a poor network; due to the slice adaptation technology, only few target slice images need to be loaded on any super-large-resolution electronic image under the current view, and even a low-end mobile phone can be easily loaded and smoothly loaded, so that the popularization and application cost of the target digital image can be greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an image processing method, apparatus, and electronic device. Background Technology

[0002] Current technologies mostly achieve high-precision replication of cultural relics and artworks through high-resolution scanning or photography, with particularly good results in the digitization of ancient calligraphy and paintings. However, significant shortcomings remain in display and detailed interaction. Specifically, existing technologies require downloading large image files (the resolution of original digital replicas of ancient calligraphy and paintings ranges from approximately 10,000×10,000 to 20,000×20,000, and the original electronic images are often 100MB-300MB in size). Download times at typical 4G internet speeds are over 10 seconds, resulting in slow file loading and cumbersome operation. Furthermore, smooth operation when loading these high-specification electronic images places considerable demands on user device performance; otherwise, stuttering, white screens, and freezes may occur, making widespread application very costly. Summary of the Invention

[0003] The purpose of this invention is to provide an image processing method, apparatus, and electronic device to reduce the cost of promoting and applying digital images.

[0004] This invention provides an image processing method, comprising: acquiring a target digital image; slicing the target digital image according to a preset slice size to obtain multiple first slice images corresponding to the target digital image; determining the number of layers matching the target digital image and the second slice images corresponding to each layer based on the TMS slicing algorithm, the preset slice size, and the multiple first slice images; receiving image operation instructions from a user for the target digital image, and determining the target layer and the target slice images in the target layer according to the image operation instructions; loading the target slice images, performing splicing and rendering processing on the target slice images to obtain a rendered image, and displaying the rendered image through a display interface.

[0005] Furthermore, the method also includes: obtaining text information corresponding to the target digital image; wherein the text information includes: image description information of the target digital image, author's biographical information, and hot data information; and converting the text information into corresponding speech information through a preset speech generation method.

[0006] Furthermore, the steps of receiving image operation instructions from the user for the target digital image and determining the target layer and the target slice image in the target layer according to the image operation instructions include: receiving image operation instructions from the user for the target digital image; determining the target layer and the slice number to be displayed according to the image operation instructions; and determining the target slice image corresponding to the slice number to be displayed from multiple slice images corresponding to the target layer according to the slice number to be displayed.

[0007] Furthermore, the method also includes: displaying multiple hotspots overlaid in the target digital image; wherein, the hotspots include: hotspot data information, image information, and voice information corresponding to the region of interest indicated by the hotspot; if a trigger command is received for a target hotspot in the displayed rendered image, displaying the hotspot data information and image information corresponding to the target region of interest indicated by the target hotspot, and / or playing the voice information corresponding to the target region of interest.

[0008] Furthermore, the method also includes: if, among at least a portion of the multiple hotspot locations, the distance between any two adjacent hotspot locations is less than a preset distance threshold, the at least a portion of the hotspot locations are aggregated to determine the aggregate location corresponding to the at least a portion of the hotspot locations; an aggregate location is displayed at the aggregate location to replace the displayed at least a portion of the hotspot locations; wherein, the aggregate location displays the attribute value corresponding to the aggregate location, and the attribute value includes at least the number of at least a portion of the hotspot locations.

[0009] Furthermore, the method also includes: generating a corresponding access link and dynamic QR code for the target digital image; if an access instruction for the access link is received from a mobile terminal, displaying the target digital image through an H5 webpage on the mobile terminal; if a scanning operation for the dynamic QR code is received, accessing the interactive digital exhibit corresponding to the target digital image.

[0010] Furthermore, the method also includes: responding to an add instruction for instructing the addition of temporary points in a target digital image; wherein the add instruction carries information to be added; the information to be added includes at least: title information to be added and description information to be added; reviewing the information to be added and obtaining a review result; if the review result indicates that the information to be added has been approved, overlaying the temporary points onto the target digital image.

[0011] This invention provides an image processing apparatus, comprising: an acquisition module for acquiring a target digital image; a slicing module for slicing the target digital image into slices of a preset slice size to obtain multiple first slice images corresponding to the target digital image; a first determination module for determining the number of layers matching the target digital image and the second slice images corresponding to each layer based on the TMS slicing algorithm, the preset slice size, and the multiple first slice images; a second determination module for receiving image operation instructions from a user for the target digital image, determining the target layer and the target slice images within the target layer according to the image operation instructions; and a display module for loading the target slice images, performing splicing and rendering processing on the target slice images to obtain a rendered image, and displaying the rendered image through a display interface.

[0012] The present invention provides an electronic device, including a processor and a memory, wherein the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the image processing method described above.

[0013] The present invention provides a machine-readable storage medium storing machine-executable instructions, which, when called and executed by a processor, cause the processor to implement any of the above-mentioned image processing methods.

[0014] This invention provides an image processing method, apparatus, and electronic device that acquires a target digital image; slices the target digital image according to a preset slice size to obtain multiple first slice images corresponding to the target digital image; determines the number of layers matching the target digital image and the second slice images corresponding to each layer based on the TMS slicing algorithm, the preset slice size, and the multiple first slice images; receives image operation instructions from the user for the target digital image, determines the target layer and the target slice images in the target layer according to the image operation instructions; loads the target slice images, performs stitching and rendering processing on the target slice images to obtain a rendered image, and displays the rendered image through a display interface. This method, through slicing technology, loads only the portion of the target slice images needed by the user each time, avoiding downloading the entire target digital image, significantly improving loading speed and response smoothness. Furthermore, by using high-sampling multi-level slicing, it maximizes the preservation of original image resolution details, ensuring a good experience even in poor network conditions. Due to the slice adaptation technology, any ultra-large resolution electronic image only needs to load a small number of target slice images in the current view, even low-end mobile phones can easily load and process them smoothly, thus greatly reducing the cost of promoting and applying target digital images. Attached Figure Description

[0015] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0016] Figure 1 A flowchart of an image processing method provided in an embodiment of the present invention; Figure 2 A schematic diagram illustrating the principle of a slicing technique provided in an embodiment of the present invention; Figure 3 This is a preview diagram of a Lushan Mountain map provided in an embodiment of the present invention; Figure 4This is a schematic diagram of a lightweight slicing process result provided in an embodiment of the present invention; Figure 5 A partial schematic diagram of a lightweight 8-level slice image provided in an embodiment of the present invention; Figure 6 A partial schematic diagram of a lightweight 5-level slice image provided in an embodiment of the present invention; Figure 7 A 256×256 pixel slice detail diagram provided for an embodiment of the present invention; Figure 8 A schematic diagram of a rendered image provided in an embodiment of the present invention; Figure 9 A schematic diagram illustrating a slice numbering rule provided in an embodiment of the present invention; Figure 10 This invention provides a custom projection specification for OpenLayers. Figure 11 A schematic diagram of slice scheduling provided in an embodiment of the present invention; Figure 12 This is a schematic diagram illustrating the loading of calligraphy and painting effects on a mobile device, as provided in an embodiment of the present invention. Figure 13 This is a schematic diagram of a ChatTTS model structure provided in an embodiment of the present invention; Figure 14 A schematic diagram of the workflow of a ChatTTS model provided in an embodiment of the present invention; Figure 15 An example of adding points to the upper part of a painting or calligraphy work according to an embodiment of the present invention; Figure 16 An example of adding dots to the lower part of a painting or calligraphy work according to an embodiment of the present invention; Figure 17 (a) A point-to-point interaction effect diagram provided by an embodiment of the present invention; Figure 17 (b) Another point-to-point interaction effect diagram provided by an embodiment of the present invention; Figure 17 (c) is another point-to-point interaction effect diagram provided by an embodiment of the present invention; Figure 17 (d) is another point-to-point interaction effect diagram provided by an embodiment of the present invention; Figure 17 (e) is another point-to-point interaction effect diagram provided by an embodiment of the present invention; Figure 17 (f) is another point-to-point interaction effect diagram provided by an embodiment of the present invention; Figure 18 An example of an aggregation effect diagram provided in an embodiment of the present invention; Figure 19 This is a progressive polymerization effect diagram provided by an embodiment of the present invention; Figure 20 A schematic diagram illustrating an annotation review mechanism provided in an embodiment of the present invention; Figure 21 This is a schematic diagram of the structure of an image processing device provided in an embodiment of the present invention; Figure 22 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0017] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Current technologies mostly achieve high-precision replication of cultural relics and artworks through high-resolution scanning or photography, with particularly good results in the digitization of ancient calligraphy and paintings. However, significant shortcomings remain in terms of display and detailed interaction. Most solutions only provide static high-definition images, and loading them requires high-end equipment and networks, making it difficult to popularize and promote them at low cost. The lack of interactive functions, limited by network and device performance, means that currently convenient mobile devices can hardly reproduce the effects of ancient calligraphy and paintings with high fidelity. These factors greatly limit the application value of ancient calligraphy and paintings in cultural tourism and education.

[0019] (2) Some ancient calligraphy and painting applications support mobile access, but most applications require users to download large files, resulting in a poor user experience due to long waiting times. In addition, existing applications generally use full-image loading, which leads to excessive device memory usage and slow response to user operations. Some low-performance machines often experience lag, white screen, and freezing, and the experience of browsing multiple images is even worse.

[0020] (3) Existing technologies are inefficient when processing large amounts of hotspot data. When the number of hotspot data exceeds one hundred, existing applications will experience lag and screen clutter. Moreover, the information contained is mostly simple text, resulting in a lack of interactive experience.

[0021] In summary, the existing technology has the following main objective drawbacks: (1) High cost of promotion and application: Existing technology requires downloading huge image files (the resolution of the original electronic reproduction of ancient calligraphy and paintings ranges from about 10,000×10,000 to 20,000×20,000, and the size of the original electronic image is often 100MB-300MB). The download time is generally more than 10 seconds at the 4G speed of the Internet. The file loading is slow and the operation is cumbersome. In addition, the smooth operation of loading the original electronic images of the above specifications requires considerable performance of the user's equipment. Otherwise, there may be lag, white screen, or freeze. The cost of promotion and application is very high.

[0022] (2) Lack of dissemination: The digital reproduction of works in the existing technology cannot be easily replicated to user devices and social media channels for sharing and experience. From the perspective of cultural tourism and teaching applications, it is impossible to enjoy the benefits of the current mobile Internet's rapid distribution and dissemination, resulting in very low brand power.

[0023] (3) Lack of interactivity: The point display in the existing technology generally does not support the interaction of hot spots of the work. It only provides an overall introduction and the author's life, and lacks professional appreciation and interpretation of the details of ancient calligraphy and painting and brushstrokes, as well as diverse materials (videos, other calligraphy and painting materials, audio materials), which fails to effectively improve the user experience.

[0024] (4) Lack of management and performance optimization of large-scale hotspots and Internet comment review mechanism: Although the existing technology supports users to submit locations, the lack of an effective management mechanism makes it difficult for users to find and display locations when there is too much data (crowding, obscuring).

[0025] Based on this, embodiments of the present invention provide an image processing method, apparatus, and electronic device, which can be applied to scenarios where it is necessary to promote and disseminate electronic versions of ancient calligraphy and paintings.

[0026] To facilitate understanding of this embodiment, an image processing method disclosed in this embodiment will first be introduced, such as... Figure 1 As shown, the method includes the following steps: Step S102: Acquire the target digital image; The aforementioned target digital image can be a high-resolution digital image obtained by scanning or photographing cultural relics, artworks, etc., using high-resolution scanning or imaging methods. In practical implementation, the target digital image can be acquired first, such as the ultra-high resolution digital image corresponding to ancient calligraphy and paintings.

[0027] Step S104: Slice the target digital image into slices according to a preset slice size to obtain multiple first slice images corresponding to the target digital image; The preset slice size can be set according to actual needs, such as dividing the slice into 256×256 pixels; in actual implementation, after obtaining the target digital image, the target digital image can be sliced ​​sequentially according to the preset slice size to obtain multiple corresponding first slice images.

[0028] Step S106: Based on the TMS slicing algorithm, the preset slice size, and multiple first slice images, determine the number of layers that match the target digital image, and the second slice image corresponding to each layer. In this embodiment, the tiling uses the TMS (Tile Map Service) tiling algorithm from geographic information technology. The TMS tiling algorithm is the core of map tile services, such as... Figure 2 The diagram illustrates the principle of a tiling technique based on a pyramid hierarchical model and a quadtree spatial index. It achieves efficient spatial data access by pre-generating multi-resolution map tiles. In the TMS algorithm, TMS tiling is based on the Web Mercator projection (EPSG:3857), converting the Earth's ellipsoid into a planar coordinate system. Map data is organized according to a multi-resolution pyramid, with each scaling level corresponding to a different resolution. Starting from the lower left corner, the data is recursively divided in the order of lower left → lower right → upper left → upper right. Each parent tile (e.g., a tile at Level n) generates four child tiles (Level n+1). Level 1 divides the globe into four tiles (2×2), and Level 2 divides it into 16 tiles (4×4). This technique is generally used to convert raw geographic data (such as Shapefiles and raster images) to the Web Mercator projection, organizing files by / level / column / row.extension, such as / 5 / 123 / 456.png.

[0029] In actual implementation, after obtaining the above multiple first slice images, the TMS slicing algorithm can be used to merge and shrink the four adjacent small images to generate multiple second slice images corresponding to each level. The size of each second slice image is the same as the preset slice size. Generally, the larger the size of the target digital image, the more levels it corresponds to.

[0030] To facilitate understanding, let's take the magnificent masterpiece "Mount Lu" by the master painter Zhang Daqian as an example. Figure 3 The image shown is a preview of the Lushan Mountain painting. The painting itself is magnificent and exquisite in detail. In order to truly reproduce its artistic charm in the digital world, we can first obtain the ultra-high resolution target digital image of the Lushan Mountain painting. Its specific parameters are: width 34252 pixels, height 6327 pixels, total pixels exceeding 216 million, and stored in JPG format.

[0031] This image contains countless delicate brushstrokes and ink tones, from the majestic overall view of Mount Lu to the details of rocks, trees, houses, and flowing springs. A close-up view of any part requires extremely high resolution to reveal the ink washes and the transitions of the brushstrokes.

[0032] The original target digital images are typically enormous, reaching hundreds of megabytes or even gigabytes in size, making it nearly impossible to load these massive images directly in web browsers or mobile applications. Loading such large amounts of data not only easily leads to network congestion and memory overflows but also causes long waiting times. Furthermore, even if the image is successfully loaded, limited by mobile phone performance, the system often experiences severe lag during interactive operations such as panning and zooming, significantly impacting the user experience. Therefore, this embodiment employs a pyramid-model-based image slicing technique for lightweight processing, enabling more efficient loading and display of this extremely high-resolution "Lushan Mountain" painting.

[0033] This solution creatively applies the TMS algorithm, originally used for slicing geographic information data, to slicing electronic files of ancient calligraphy and paintings. A Python-based slicing algorithm is developed, taking the electronic file of "Lushan Map" (i.e., the target digital image) as input. Without setting the geographic coordinates of the original image, the original image (i.e., the target digital image) is divided into 256×256 pixel relative coordinates, and then sampled and sliced. The actual coordinates of each slice image are calculated sequentially. Gdal is used to slice the original image in a sampling-image generation-numbering-storage manner.

[0034] like Figure 4 The diagram shows a lightweight tile processing result. This method can automatically generate a TMS standard metadata file tilemapresource.xml, injecting information such as tile level and resolution, which is convenient for use by the geographic information engine OpenLayers.js.

[0035] In this embodiment, for the target digital image corresponding to the "Lushan Map", the original image is first sliced ​​into a maximum of 8 levels according to the maximum electronic image resolution. Each slice is then divided into 256×256 tiles, covering the entire original image. Next, Gdal is used to sample the 8-level slices, and tiles from level 7 to level 0 are generated sequentially based on the TMS slicing algorithm. For example... Figure 5 A partial schematic diagram of a lightweight 8-level slice image is shown, and as follows: Figure 6 This is a partial schematic diagram of a lightweight 5-level slice image. (See attached image.) Figure 7 The image shown is a detailed schematic diagram of a 256×256 pixel slice.

[0036] Finally, after slicing, the target digital image is divided into multiple small blocks, each a separate slice file. Each level of slice image has a different resolution, corresponding to different scaling levels; the higher the level, the higher the resolution. For example, if the level range is 0-8, level 8 has the highest resolution, and level 0 has the lowest. The number and resolution of slice images are dynamically adjusted according to the scaling level. The size of each slice image is optimized to preserve detail without excessively occupying storage space. In this embodiment, based on actual calligraphy and painting data, the highest level of slicing is level 8. Because the smooth loading of different levels of slice images needs to be dynamically scheduled according to user operations, the actual total information content of the slice images is increased by 2.5 times compared to the original image, which is a space-for-time trade-off.

[0037] Step S108: Receive the user's image operation command for the target digital image, and determine the target layer and the target slice image in the target layer according to the image operation command; The aforementioned image operation instructions may be movement instructions, zoom-out instructions, zoom-in instructions, etc.; the aforementioned target level may be any level among multiple levels; and the aforementioned target slice image may be at least a portion of the multiple slice images contained in the target level.

[0038] Step S110: Load the target slice image, perform stitching and rendering processing on the target slice image to obtain a rendered image, and display the rendered image through the display interface.

[0039] Slicing technology makes image display more flexible and efficient. Users can freely zoom and browse the "Lushan Mountain Map" according to their needs. Whether viewing the grand panorama of the entire map or zooming in to examine the fine brushstrokes and ink layers of a specific section, the system will automatically load the appropriate target slice image based on the current zoom level. (According to the TMS algorithm principle, slices can be called at any position and level. The current view on the screen consists of a maximum of 16 slice images. The size and number of 256×256 pixel images of each slice at the current level are controllable.) Figure 8 The diagram illustrates a rendering method that supports smooth access even with low network and device performance. It ensures image quality and display while avoiding stuttering and loading delays. The sliced ​​images not only significantly improve loading speed but also provide the optimal browsing experience based on device performance. By breaking down GB-level image loads into numerous KB-level 256×256 pixel slices in smaller requests, with a maximum of 16 slices covering the current view, network transmission pressure and client memory usage are greatly reduced. This ensures second-level initial loading and smooth subsequent interactions even in mobile network environments with limited bandwidth.

[0040] The image processing method described above involves: acquiring a target digital image; slicing the target digital image according to a preset slice size to obtain multiple first slice images corresponding to the target digital image; determining the number of layers matching the target digital image and the second slice images corresponding to each layer based on the TMS slicing algorithm, the preset slice size, and the multiple first slice images; receiving user image operation commands for the target digital image; determining the target layer and the target slice images within the target layer based on the image operation commands; loading the target slice images; performing stitching and rendering processing on the target slice images to obtain a rendered image; and displaying the rendered image through a display interface. This method, through slicing technology, loads only the portion of the target slice images needed by the user each time, avoiding downloading the entire target digital image, significantly improving loading speed and responsiveness. Furthermore, by using high-sampling, multi-level slicing, it maximizes the preservation of original image resolution details, ensuring a good experience even in poor network conditions. Due to the slice adaptation technology, any ultra-large resolution electronic image only requires loading a small number of target slice images in the current view, allowing even low-end mobile phones to easily load and process them smoothly, thus greatly reducing the cost of promoting and applying target digital images.

[0041] This invention also provides another image processing method, which is implemented based on the method in the above embodiments, and includes the following steps: Step 1: Acquire the target digital image; Step 2: Slice the target digital image into slices according to the preset slice size to obtain multiple first slice images corresponding to the target digital image; Step 3: Based on the TMS slicing algorithm, the preset slice size, and multiple first slice images, determine the number of layers that match the target digital image, and the second slice image corresponding to each layer. Step 4: Receive the user's image operation instructions for the target digital image, and determine the target layer and the slice number to be displayed based on the image operation instructions; Step 5: Based on the slice number to be displayed, determine the target slice image corresponding to the slice number to be displayed from multiple slice images corresponding to the target level.

[0042] In this embodiment, Gdal can automatically generate a corresponding number for each generated slice image, for example, such as... Figure 9 The diagram illustrates a slice numbering rule. After receiving an image operation command from a user, the corresponding target level and the slice number to be displayed in that target level can be determined based on the image operation command; then, the corresponding target slice image can be determined based on the slice number to be displayed.

[0043] Step 6: Load the target slice image, perform stitching and rendering on the target slice image to obtain the rendered image, and display the rendered image through the display interface.

[0044] To more efficiently present and interactively view ultra-high-resolution images like the "Lushan Map," lightweight data processing (such as the aforementioned tiling technique) is necessary, as is optimizing front-end loading and the user experience. Therefore, this embodiment employs tile rendering technology based on a map engine to achieve a smoother mobile loading and interactive experience.

[0045] In this solution, the system frontend abandons traditional image viewers and instead integrates a mature open-source map rendering engine (OpenLayers). This engine treats the pre-processed tile images as a custom "map layer." The engine's core scheduler listens for user interaction events (such as dragging, zooming, and rotating), calculates the current view state in real time, and sends asynchronous requests to the tile server to obtain the corresponding set of tile URLs. The engine then precisely stitches and renders these tiles (corresponding to the target tile images mentioned above) within a Canvas or WebGL context to form a complete visual image.

[0046] The core logic of graphics loading and scheduling is explained below: Since the objects of the slices are not TIF (Tagged Image File Format, a general raster image storage format) geographic information images with coordinates, the target slice images cannot be loaded directly using the coordinate system preset by OpenLayers or any existing geographic coordinate system. This is because ancient calligraphy and painting slices do not have any geographic information attributes, and direct projection onto the map will result in loading errors.

[0047] Map projection is a mathematical method for representing the three-dimensional surface of the Earth on a two-dimensional plane. OpenLayers supports two common projections by default: EPSG:4326 (WGS84: World Geodetic System 1984): a geographic coordinate system that uses latitude and longitude to represent location; and EPSG:3857 (Web Mercator): a standard projection for web maps, with units in meters. When the data source is based on a coordinate system other than the two mentioned above, a custom projection is required. Common scenarios include using local coordinate systems, area-specific projections (such as the area-specific conic projection for China), UTM (Universal Transverse Mercator) regional projections, or polar projections. This solution uses a custom projection: a planar projection that projects a two-dimensional image (target digital image) onto a map.

[0048] This solution innovatively employs a custom projection method to load non-geographic image tiles during graphics loading and scheduling. Custom projection is a method that abandons geographic information projection. It uses a planar direct coordinate direct projection method, defining the width and height of the entire projection canvas (the original width and height of the ancient calligraphy and painting electronic file), defining the origin of the TMS tile, and then loading the tile in the traditional OpenLayers loading method for TMS layers.

[0049] OpenLayers custom projection example code: const customProjection = new ol.proj.Projection({ code: 'CUSTOM', / / Code name for the custom projection units: 'm', / / Unit, usually 'degrees' or 'm' (meter) extent: [0, -height, width, 0], / / The range of the projection, which in this scheme is a rectangle bounded by coordinates 0, 0 to the width and height of the original electronic component image. }); like Figure 10 The illustration shows a custom projection method in OpenLayers, which determines a target slice image based on a target digital image, tiles the target slice image according to the corresponding grid, and finally forms a rendered image.

[0050] This method leverages OpenLayers' ability to directly load TMS-compliant tiles as a geographic information system, automatically calculating the actual tile number based on the user's current screen coordinates and scheduling the loading of the corresponding tile in the format {z} / {x} / {-y}.png to complete rendering. It also employs custom projection to avoid the problem of tiles without projection, by setting the tile origin, projection canvas width, and height to always lock the tile position to the center of the screen canvas.

[0051] like Figure 11 The diagram illustrates a tile scheduling approach. The current view can display a maximum of 16 tile images, with the grid numbers representing the tile numbers. When the screen is in this position, OpenLayers calculates the canvas coordinates based on the current screen coordinates, obtaining the layer and corresponding tile number of each of the 16 tile images. Since the 16 tile images belong to the same layer, a resource path is constructed based on the layer, the corresponding tile number, and the current zoom level (e.g., 3 / 1 / 1.png). The corresponding target tile image is then directly requested from the server and rendered. Other target tile images can be processed in the same way. When the screen moves, a new tile layer and tile number are calculated, and the same method is used for requesting and rendering. The previous batch of tiles can then be removed from memory, saving computing resources.

[0052] like Figure 12 The illustration shows a mobile application loading effect for calligraphy and painting. This method provides an experience almost identical to that of a mobile map application, ensuring users can perform smooth operations such as two-finger zooming, swiping panning, and inertial scrolling, quickly and without lag browsing every part of a large-scale painting. Through layered slicing technology, loading and display are dynamically based on the user's view range, avoiding redundant data transmission and improving loading speed and responsiveness. Users do not need to wait for the entire image to load; they can freely explore the panorama or details at any time, experiencing a smooth and immediate operation that greatly enhances interactivity and fluidity.

[0053] This approach leverages deep optimizations in graphics rendering and gesture interaction at the map engine's underlying level, providing smoother operation feedback far exceeding traditional webpage image viewers. It can achieve real-time operation at the maximum frame rate of the current device's browser. Developed based on the HTML5 (Hypertext Markup Language 5) standard, it can run seamlessly in browsers on various mobile devices such as iOS and Android without requiring the installation of native applications, achieving the cross-platform advantage of "develop once, run anywhere".

[0054] Step 7: Obtain the text information corresponding to the target digital image; the text information includes: image description information of the target digital image, author's biographical information, and hot data information; In this embodiment, in addition to the target digital image, it usually also includes a large amount of detailed content corresponding to the target digital image. This content revolves around the target digital image and mainly includes: image description information of the target digital image, biographical information of the author, multiple hot spots of the painting and hot spot data information of each hot spot. This information usually requires the participation and improvement of art appreciation experts or the use of LLM (Large Language Model) to retrieve and improve online information to form a set of auxiliary explanatory materials for the work, that is, the text information corresponding to the target digital image.

[0055] Step 8: Convert the text information into corresponding speech information using a preset speech generation method.

[0056] In this embodiment, generative AI (Artificial Intelligence) technology can be used to directly generate speech from the above text information, saving a lot of manpower and time costs.

[0057] The specific technology is as follows: It is implemented using the ChatTTS open-source speech generation system. It adopts an LLM-driven context-aware generation framework, builds a dialogue understanding module based on the Transformer-XL or GPT-4 architecture, captures long-distance dialogue dependencies through the Local+GlobalAttention mechanism, supports a context window of 512+ tokens, adopts the Dual-PathEncoding scheme to independently encode the current question (Query) and the history of the dialogue, and then performs semantic alignment through the Cross-Attention layer. It introduces a Belief State Tracker module to dynamically maintain meta-information such as dialogue entities, intents, and sentiment states, and generates hidden state vectors containing the dialogue history.

[0058] like Figure 13 The diagram shown illustrates the ChatTTS model structure. ChatTTS's core design philosophy is to employ an end-to-end generative model, integrating multiple independent steps in traditional speech synthesis (such as text processing, acoustic feature generation, and waveform synthesis) into a unified framework. Its backbone network is an autoregressive generative model based on the Transformer architecture, specifically optimized for dialogue scenarios. The core design of the overall architecture can be viewed as a system that converts text sequences into audio token sequences and then reconstructs them into waveforms (audio). For details on the model structure, please refer to relevant technologies; further elaboration is omitted here.

[0059] like Figure 14 The diagram illustrates the workflow of a ChatTTS model. After receiving text input, the ChatTTS model preprocesses the input text, such as performing text normalization and word segmentation, and then performs prosodic prediction. Specifically, it analyzes the text and predicts prosodic features such as pauses and stresses. After that, it generates a token sequence, which can be predicted using the LlaMA Transformer model to predict the corresponding multi-scale audio token sequence. Then, it performs waveform reconstruction and finally outputs the audio. For details, please refer to relevant technologies, which will not be elaborated here.

[0060] In this embodiment, the out-of-vocabulary (OOV) problem can be solved by using a BERT-based word segmenter at the front end, combined with a domain dictionary (such as a medical / financial terminology database). Multi-task learning is used to jointly predict phoneme duration, fundamental frequency (F0) contour, and energy profile, and a Transformer-Decoder structure is used to model prosodic sequences. At the back end, acoustic generation is based on a diffusion model. Multi-band Mel-Spectrogram Prediction enhances the ability to restore high-frequency details, and causal convolution and adaptive instance normalization (AdaIN) are employed to achieve 48kHz high-fidelity audio output.

[0061] The training method of the ChatTTS model is explained below. The model is unsupervised pre-trained on 1 million hours of publicly available speech data. For example, the publicly available speech data can be LibriTTS (A Corpus Derived from LibriSpeech for Text-to-Speech) or VCTK (English Multi-speaker Corpus for CSTR Voice Cloning Toolkit; CSTR is an abbreviation for "Centre for Speech Technology Research"). Fine-tuning is performed using 50,000 hours of vertical domain dialogue data (such as customer service / navigation) to optimize the pronunciation accuracy of domain-specific words (such as product models and road names). It is one of the best frameworks in the field of generative AI text-to-graph in the current open source field. By generating diverse multi-turn dialogue data through back-translation and role-playing to simulate real-world conversation environments, background noise (such as café / car ambient sounds) and channel distortion are added. Wasserstein GAN is introduced to improve the naturalness of the generated speech and enhance the robustness of the application.

[0062] This model addresses the problems of traditional TTS (Text-to-Speech, the traditional method is speech synthesis), such as fragmented context, monotonous emotion, and insufficient real-time performance.

[0063] This approach uses LLM large model retrieval combined with art appreciation expert annotation to produce content accompanying electronic ancient calligraphy and paintings. It also uses the generative AI framework ChatTTS to further generate text and speech content. For example, in this case, speech was generated for more than 20 hot texts. All relevant results will be applied to the subsequent step - the interactive step of adding location information to the image.

[0064] Step 9: Multiple hotspots are overlaid in the target digital image; each hotspot contains: hotspot data information, image information, and voice information corresponding to the region of interest indicated by the hotspot. The aforementioned region of interest can be understood as the location of hotspots in the target digital image. For example, if the target digital image includes a pavilion, a hotspot can be added at the location of the pavilion, containing hotspot data, image information, and audio information related to the pavilion. In this embodiment, OpenLayers provides an efficient way to add points of interest to an image. This function uses the vector layer function of the map engine, combined with geographic coordinate registration technology, to accurately overlay key points of interest from artworks onto the target digital image, thereby achieving deep interaction between cultural tourism and education.

[0065] The following explains the principle of point feature addition: OpenLayers uses vector layers to display hotspots on top of the layer corresponding to the sliced ​​image. Hotspots are overlaid on the layer corresponding to the sliced ​​image, and they do not intersect with the image. Each hotspot represents a key element in the painting (such as landscape layout, figure posture, inscriptions, seals, specific brushstrokes, etc.). These vector point features (i.e., hotspots) are precisely overlaid at their corresponding positions in the painting using geographic coordinate registration technology. For details on the principle of point feature addition, please refer to relevant technical documentation; it will not be elaborated upon here. Each hotspot feature is defined as a data object, typically containing the following main contents: Canvas coordinates: Strictly speaking, a map engine should normally display latitude and longitude coordinates, but since this solution uses custom canvas coordinates, which are actually the width and height of a custom canvas, corresponding to the XY position in the entire target digital image, they are used to accurately locate hotspots on the corresponding positions on the canvas. The aforementioned hotspot data information needs to be filled into the hotspot position.

[0066] Attribute set: Each hotspot location usually also contains an attribute set, which stores more information about the hotspot location, such as text descriptions, image information, audio narration, etc., which can be used to enrich the user's interactive experience.

[0067] Example of location data format: { "points": [ { "id": "p001", / / Unique identifier for the location "canvasCoord": { / / Custom canvas coordinates (corresponding to the XY position of the image) "x": 40605.891694880287, / / Horizontal coordinate (in pixels, based on the width and height of the digitally reproduced image of "Lushan Map") "y": -64931.52947890446 / / Vertical coordinate (in pixels) }, "properties": { / / Property set: stores detailed information about the location "name": "Artwork Appreciation 1", / / Location Name "type": "Landscape Layout", / / Location type (e.g., landscape, figures, seals, etc.) "description": "The beginning section depicts misty clouds, peaks appearing and disappearing, and smoke rising from the valley floor." / / Text description "image": "assets / images / p001_detail.jpg", / / Path to the detail image "audio": "assets / audios / p001_narration.mp3", / / Audio narration path "audioText": "The main peak waterfall is one of the visual focal points of the 'Lushan Mountain Map'...", / / Text corresponding to the audio (used for pre-generated audio) "author": "Narrator: Professor Li (China Academy of Art)", / / Narration Information "timestamp": "2023-10-01" / / Information update time } } ] } Hotspots added in this way have the characteristic of scaling without distortion. The underlying calculation principle is: Each point is bound to a screen coordinate corresponding to its XY position on the canvas, and the point is rendered at that location. Whenever the map zooms, causing the previous screen coordinate position to change, the new screen coordinates corresponding to the XY position need to be recalculated, and the point is rendered again at the new screen coordinates. This process is repeated to ensure that all points always move and render in the correct position on the canvas during dynamic zooming and translation.

[0068] In this way, important parts of the "Lushan Mountain Map" (such as figures, landscapes, and seals) can be presented as interactive points, allowing users to view detailed information about each point of interest by clicking or swiping. This interactive experience not only enhances the visualization of the artwork but also effectively supports cultural tourism and educational applications.

[0069] Step 10: If a trigger command is received for a target hotspot in the displayed rendered image, display the hotspot data information and image information corresponding to the target region of interest indicated by the target hotspot, and / or play the voice information corresponding to the target region of interest.

[0070] like Figure 15 The image shown illustrates the effect of adding points to the upper part of a painting or calligraphy work. Figure 16 This is an example of adding effect points to the lower part of a painting or calligraphy work. After clicking on the target hotspot point, OpenLayers reads the point's attribute set, retrieves relevant data, and performs rendering. Clicking the speaker icon can play pre-generated audio corresponding to the text. (Example...) Figure 17 (a) shows a point-to-point interaction effect diagram. Figure 17 (b) shows another point-to-point interaction effect diagram. Figure 17 (c) shows another point-to-point interaction effect diagram. Figure 17 (d) shows another point-to-point interaction effect diagram. Figure 17 (e) shows another point-to-point interaction effect diagram. Figure 17 (f) shows another point-to-point interaction effect diagram.

[0071] Step 11: If, among at least some of the multiple hot spots, the distance between any two adjacent hot spots is less than a preset distance threshold, aggregate the at least some hot spots to determine the aggregated location corresponding to the at least some hot spots. Step 12: Display an aggregation point at the aggregation location to replace the display of at least some hotspot points; wherein, the aggregation point displays the attribute value corresponding to the aggregation point, and the attribute value includes at least the number of at least some hotspot points.

[0072] The aforementioned preset distance threshold can be set according to actual needs, and is usually a relatively small distance value. In actual implementation, when the number of hotspots is too large, even affecting the unfolding of the artwork, point aggregation technology can be used. Point aggregation is implemented through ol.source.Cluster. The principle of aggregation is to merge hotspots that are relatively close into a single aggregated point and calculate the attribute values ​​of the merged aggregated point. In the aggregation source ol.source.Cluster, when a new hotspot is added, it checks whether the distance between the new hotspot and existing aggregated points is within the specified aggregation distance. If so, the new hotspot is added to the aggregated point, and the attribute values ​​of the aggregated point (such as the number of aggregated hotspots) are updated. If the distance between the new hotspot and existing aggregated points exceeds the specified aggregation distance, the new hotspot is added to the aggregation source as a new aggregated point. In this embodiment, the location of an existing aggregated point can be the center of the area formed by multiple aggregated hotspots. The distance between the new hotspot and this center location can be calculated and compared with the specified aggregation distance. During rendering, different styles can be set for aggregated points based on their attribute values ​​to distinguish them from ordinary hotspot points. For example... Figure 18 The image shown is an example of an aggregation effect, and Figure 19 This image shows a progressive aggregation effect. The numbers circled represent the number of hotspots aggregated at that aggregation point. As the image is zoomed in, the hotspots within the aggregation point can be displayed sequentially.

[0073] In this embodiment, the reasonable density of information presentation ensures both a macro-level overview and the preservation of all micro-level details, guiding users through exploratory learning. This approach achieves a one-to-one spatial correspondence between image content and interpretive information, firmly binding abstract descriptive text with concrete visual elements, greatly improving the accuracy and efficiency of knowledge transfer. Each hotspot is a structured data entity, providing a solid data foundation for subsequent aggregation analysis, filtering queries, and multimedia expansion. When processing a large number of hotspots, especially at smaller zoom levels (i.e., when the map display area is wide), multiple hotspots may overlap due to excessive spatial density, causing "visual pollution." This not only affects the user's visual experience but may also make it difficult for users to understand or manipulate images. To solve this problem, hotspots can be effectively aggregated and distributed to ensure a clear and readable interface at different zoom levels.

[0074] Step 13: Generate a corresponding access link and dynamic QR code for the target digital image; Step fourteen: If an access instruction for the access link is received from the mobile terminal, the target digital image is displayed on the H5 webpage of the mobile terminal. Step 15: If a scanning operation for the dynamic QR code is received, proceed to the interactive digital exhibit corresponding to the target digital image.

[0075] The display technology for digitally reproduced works like Zhang Daqian's "Lushan Mountain Map" suffers from poor dissemination in the cultural tourism and education sectors. Particularly in these areas, it fails to fully leverage the rapid distribution advantages of the current mobile internet, resulting in extremely limited brand communication. Users cannot easily share the corresponding target digital images to personal devices or self-media platforms during the experience, thus affecting the breadth and depth of its dissemination. To address this issue, this embodiment generates a permanent and unique access link (URL: Uniform Resource Locator) and a bound dynamic QR code for each independent target digital image. The link points to a mobile-optimized H5 webpage. These dynamic QR codes can be embedded in offline materials such as museum exhibit labels, promotional posters, educational manuals, academic publications, and tickets. Users can simply scan the code with any QR code scanning application on their mobile device while visiting the site or reading materials to immediately access the corresponding interactive digital exhibit, without any cumbersome download, installation, or search process.

[0076] Step sixteen: Respond to the add instruction used to instruct the addition of temporary points in the target digital image; wherein the add instruction carries information to be added; the information to be added includes at least: title information to be added and description information to be added; Step 17: Review the information to be added and obtain the review results; Step 18: If the review result indicates that the information to be added has been approved, overlay the temporary location onto the target digital image.

[0077] In large-scale site displays, information retrieval and utilization often face problems of low efficiency and insufficient accuracy, especially in scenarios requiring dynamic updates and multi-party participation. To ensure the quality of generated content, this solution introduces an internet-based annotation and review mechanism. Authorized users (such as registered art enthusiasts and researchers) can use a dedicated tool on the front-end interface to "pin" temporary points at designated locations on the artwork, filling in relevant information to be added, such as titles and descriptions, and uploading supporting images or videos. These submissions first enter the back-end management system for review and fact-checking to ensure authenticity and high quality. In practice, this review process can also be performed manually. Once approved, the temporary point is officially published, becoming part of the main site data pool and publicly visible. The temporary point can be overlaid onto the target digital image and made available for all users to search and view, thereby enhancing the interactivity and credibility of the information.

[0078] like Figure 20 The diagram illustrates an annotation review mechanism. Based on user-submitted annotations, temporary data points are generated. The submitted content is then reviewed. If the review is successful, the temporary data point can be converted into a formal data point and included in the data pool for public access. If the review fails, the user is notified of the reason and allowed to modify and resubmit the data.

[0079] This embodiment, through this innovative approach, not only enhances user engagement, enabling them to easily access interactive digital exhibits without any complex operations, but also lowers the technical barrier through dynamic QR codes, making it suitable for users of different ages and technical backgrounds.

[0080] Meanwhile, the expanded exhibit placement feature incorporates a crowdsourcing mechanism, enriching the interpretation of exhibits and ensuring the professionalism and accuracy of the content during the review process, thus safeguarding the platform's authority and academic value. Ultimately, this mechanism brings greater efficiency, better quality control, and broader social participation to the dissemination of digital culture, forming a sustainable digital cultural ecosystem.

[0081] Traditional online display methods typically involve loading the entire image directly or loading it in simple chunks, which has inherent drawbacks such as slow initial loading, limited interaction modes, inability to carry multi-dimensional information, difficulty in internet dissemination, and lack of user engagement.

[0082] This solution is a comprehensive cultural tourism and education solution that uses ultra-high resolution digital images, such as electronic replicas of ancient calligraphy and paintings, as the original material. Through complex technology, it achieves high fidelity while being lightweight enough for mobile devices with low network and hardware requirements. It utilizes the internet for dissemination, uses generative AI to generate audio content, and involves experts in constructing the core content framework and establishing an internet review mechanism to create a community for the public to co-build a cultural and artistic community.

[0083] This innovative solution combines modern web mapping technology, lightweight image processing technology, generative AI, and a user-co-built cultural artifact community review mechanism to construct a complete technical solution encompassing data preprocessing, mobile interaction, internet social media distribution, and community operation. Through image slicing and tile loading, map engine-driven architecture, intelligent location management, and multimedia expansion, it ultimately achieves lightweight access, immersive interaction, intelligent interpretation, and sustainable ecological operation of cultural heritage on mobile devices. This significantly enhances its application value and public interest participation in education, exhibitions, tourism, and research, driving traffic to related fields and providing a revolutionary technical solution for establishing an online digital cultural artifact community.

[0084] The image processing method described above, through slicing technology, loads only the portion of data needed by the user at a time, avoiding the download of large files and significantly improving loading speed and responsiveness. Furthermore, by using high-sampling, multi-level slicing, it preserves the original image resolution details to the maximum extent, ensuring a good experience even in environments with poor network conditions. Due to the slicing adaptation technology, any ultra-large resolution electronic image only requires loading a dozen or so small images in the current view, making it easy to load and process even on low-end mobile phones with smooth loading. In specific applications, venues can directly provide QR codes for users to scan and use without preparing dedicated display equipment, thus reducing application costs.

[0085] In this method, users can browse, share, and comment on ancient calligraphy and paintings through mobile H5 web pages, and QR codes can be embedded directly in the venue and social media to facilitate users to scan and view and forward them directly in WeChat, which greatly enhances the dissemination and topicality of ancient calligraphy and paintings in the cultural tourism and education industries.

[0086] This solution introduces a hotspot display feature, which supports interaction and voice narration. After a user clicks on a specific location, background information, historical details, and artistic interpretations can be provided via voice, helping the user understand the historical and cultural connotations of the artwork more vividly and intuitively. Furthermore, it supports viewing detailed hotspot information in text format, allowing users to explore the information of each location more comprehensively. To enhance usability, users can easily and quickly locate and browse artworks of interest in the virtual space, improving immersion, interactivity, and the overall user experience.

[0087] This approach also introduces point aggregation, allowing hundreds of points to be displayed together with excellent performance. The display method is dynamically adjusted according to the scaling level to avoid interface clutter. This approach can also introduce search technology, allowing users to search for hotspots by keywords. At the same time, through an approval mechanism, the high quality of annotations added by authorized users can be guaranteed, thereby ensuring the quality and management efficiency of newly added temporary point data. This makes the display of large-scale points more orderly, and users can find and use information more conveniently.

[0088] This invention provides an image processing apparatus, such as... Figure 21 As shown, the device includes: an acquisition module 220 for acquiring a target digital image; a slicing module 221 for slicing the target digital image into slices according to a preset slice size to obtain multiple first slice images corresponding to the target digital image; a first determination module 222 for determining the number of layers matching the target digital image and the second slice image corresponding to each layer based on the TMS slicing algorithm, the preset slice size, and the multiple first slice images; a second determination module 223 for receiving image operation instructions from the user for the target digital image, determining the target layer and the target slice images in the target layer according to the image operation instructions; and a display module 224 for loading the target slice images, performing splicing and rendering processing on the target slice images to obtain a rendered image, and displaying the rendered image through a display interface.

[0089] The aforementioned image processing device, through slicing technology, loads only the portion of the target image that the user needs at a time, avoiding the download of the entire target digital image. This significantly improves loading speed and responsiveness. Furthermore, by using high-sampling, multi-level slicing, it maximizes the preservation of original image resolution details, ensuring a good experience even in environments with poor network conditions. Due to the slicing adaptation technology, any ultra-large resolution electronic image only requires loading a small number of target slice images in the current view. Even low-end mobile phones can easily load and process the image smoothly, thereby greatly reducing the cost of promoting and applying target digital images.

[0090] Furthermore, the device is also used to: acquire text information corresponding to the target digital image; wherein the text information includes: image description information of the target digital image, author's biographical information, and hot data information; and convert the text information into corresponding speech information through a preset speech generation method.

[0091] Furthermore, the second determining module is also used to: receive image operation instructions from the user for the target digital image; determine the target level and the slice number to be displayed based on the image operation instructions; and determine the target slice image corresponding to the slice number to be displayed from multiple slice images corresponding to the target level based on the slice number to be displayed.

[0092] Furthermore, the device is also used to: overlay multiple hotspots in a target digital image; wherein the hotspots include: hotspot data information, image information, and voice information corresponding to the region of interest indicated by the hotspot; if a trigger command is received for a target hotspot in the displayed rendered image, the device displays the hotspot data information and image information corresponding to the target region of interest indicated by the target hotspot, and / or plays the voice information corresponding to the target region of interest.

[0093] Furthermore, the device is also used to: if, among at least a portion of the multiple hotspot locations, the distance between any two adjacent hotspot locations is less than a preset distance threshold, aggregate the at least a portion of the hotspot locations to determine the aggregate location corresponding to the at least a portion of the hotspot locations; display an aggregate location at the aggregate location to replace the display of the at least a portion of the hotspot locations; wherein, the aggregate location displays an attribute value corresponding to the aggregate location, and the attribute value includes at least the number of the at least a portion of the hotspot locations.

[0094] Furthermore, the device is also used to: generate corresponding access links and dynamic QR codes for the target digital image; if an access instruction for the access link is received from a mobile terminal, display the target digital image through the H5 webpage of the mobile terminal; if a scanning operation for the dynamic QR code is received, access the interactive digital exhibit corresponding to the target digital image.

[0095] Furthermore, the device is also used to: respond to an add instruction for instructing the addition of temporary points in a target digital image; wherein the add instruction carries information to be added; the information to be added includes at least: title information to be added and description information to be added; review the information to be added and obtain a review result; if the review result indicates that the information to be added has been approved, overlay the temporary points onto the target digital image.

[0096] The image processing apparatus provided in this embodiment of the invention has the same implementation principle and technical effect as the aforementioned image processing method embodiment. For the sake of brevity, any parts not mentioned in the image processing apparatus embodiment can be referred to the corresponding content in the aforementioned image processing method embodiment.

[0097] This invention also provides an electronic device, see [link to relevant documentation]. Figure 22 As shown, the electronic device includes a processor 130 and a memory 131. The memory 131 stores machine-executable instructions that can be executed by the processor 130. The processor 130 executes the machine-executable instructions to implement the above-described image processing method.

[0098] Furthermore, Figure 22 The electronic device shown also includes a bus 132 and a communication interface 133, with the processor 130, the communication interface 133 and the memory 131 connected via the bus 132.

[0099] The memory 131 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 133 (which can be wired or wireless), such as the Internet, wide area network, local area network, metropolitan area network, etc. The bus 132 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 22 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.

[0100] Processor 130 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 130 or by instructions in software form. Processor 130 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 131. The processor 130 reads the information from memory 131 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.

[0101] This invention also provides a machine-readable storage medium storing machine-executable instructions. When these machine-executable instructions are called and executed by a processor, they cause the processor to implement the above-described image processing method. For specific implementation details, please refer to the method embodiments, which will not be repeated here.

[0102] The computer program products of the image processing method, apparatus, and electronic device provided in the embodiments of the present invention include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.

[0103] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0104] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An image processing method, characterized in that, The method includes: Acquire the target digital image; The target digital image is sliced ​​according to a preset slice size to obtain multiple first slice images corresponding to the target digital image; Based on the TMS slicing algorithm, the preset slice size, and multiple first slice images, the number of layers matching the target digital image and the second slice image corresponding to each layer are determined. Receive image operation instructions from the user for the target digital image, and determine the target layer and the target slice image in the target layer according to the image operation instructions; The target slice image is loaded, and the target slice image is stitched and rendered to obtain a rendered image, which is then displayed through a display interface.

2. The method according to claim 1, characterized in that, The method further includes: Obtain the text information corresponding to the target digital image; wherein, the text information includes: image description information of the target digital image, author's biographical information, and hot data information; The text information is converted into corresponding voice information using a preset voice generation method.

3. The method according to claim 1, characterized in that, The steps of receiving image operation instructions from a user for the target digital image, and determining the target layer and the target slice image within the target layer based on the image operation instructions, include: Receive image operation instructions from the user for the target digital image, and determine the target level and the slice number to be displayed based on the image operation instructions; Based on the slice number to be displayed, the target slice image corresponding to the slice number to be displayed is determined from multiple slice images corresponding to the target level.

4. The method according to claim 1, characterized in that, The method further includes: The target digital image displays multiple hotspots overlaid on it; wherein, each hotspot contains: hotspot data information, image information, and voice information corresponding to the region of interest indicated by the hotspot; If a trigger command is received for a target hotspot in the displayed rendered image, the hotspot data information and image information corresponding to the target region of interest indicated by the target hotspot are displayed, and / or the voice information corresponding to the target region of interest is played.

5. The method according to claim 4, characterized in that, The method further includes: If, among at least a portion of the multiple hotspot locations, the distance between any two adjacent hotspot locations is less than a preset distance threshold, the at least a portion of the hotspot locations are aggregated to determine the aggregated location corresponding to the at least a portion of the hotspot locations. An aggregation point is displayed at the aggregation location to replace the display of the at least some hotspot points; wherein, an attribute value corresponding to the aggregation point is displayed at the aggregation point, and the attribute value includes at least the number of the at least some hotspot points.

6. The method according to claim 1, characterized in that, The method further includes: Generate a corresponding access link and dynamic QR code for the target digital image; If an access instruction for the access link is received from a mobile terminal, the target digital image is displayed on the H5 webpage of the mobile terminal; If a scan operation is received for the dynamic QR code, the user will be directed to the interactive digital exhibit corresponding to the target digital image.

7. The method according to claim 1, characterized in that, The method further includes: The response is an instruction to add temporary points to the target digital image; wherein the addition instruction carries information to be added; the information to be added includes at least: title information to be added and description information to be added; The information to be added is reviewed, and the review result is obtained; If the review result indicates that the information to be added has passed the review, the temporary location will be overlaid onto the target digital image.

8. An image processing apparatus, characterized in that, The device includes: The acquisition module is used to acquire the target digital image; The slicing module is used to slice the target digital image according to a preset slice size to obtain multiple first slice images corresponding to the target digital image; The first determining module is used to determine the number of layers that match the target digital image and the second slice image corresponding to each layer based on the TMS slicing algorithm, the preset slice size and multiple first slice images; The second determining module is used to receive image operation instructions from the user for the target digital image, and determine the target layer and the target slice image in the target layer according to the image operation instructions; The display module is used to load the target slice image, perform splicing and rendering processing on the target slice image to obtain a rendered image, and display the rendered image through the display interface.

9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, the processor executing the machine-executable instructions to implement the image processing method according to any one of claims 1-7.

10. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions that, when invoked and executed by a processor, cause the processor to implement the image processing method according to any one of claims 1-7.