Image display method and apparatus therefor
By displaying images in a virtual 3D space and using multimodal features and metadata to determine camera pose, the problem of the single image resource display method in the existing technology is solved, and efficient, intuitive and interesting 3D image display is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 艾酷软件技术(上海)有限公司
- Filing Date
- 2026-02-12
- Publication Date
- 2026-06-02
AI Technical Summary
Existing image management and display technologies lack spatial depth in their display methods, making it difficult to manage and display image resources efficiently, intuitively, and in an engaging way.
By receiving input, extracting multimodal features and metadata of the image, determining the camera pose, and displaying the image in a virtual 3D space, the image is displayed in a virtual 3D space using multimodal features and metadata, and combined with physical rendering and depth estimation techniques to achieve 3D image display.
It achieves efficient, intuitive, and engaging 3D image display, enhancing the user experience.
Smart Images

Figure CN122134984A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of electronic equipment technology, and specifically relates to an image display method and apparatus. Background Technology
[0002] With the rapid development of mobile internet and smart terminal technologies, smartphones, tablets, and wearable devices equipped with high-pixel camera modules have become ubiquitous, and photography has become an important way for users to record their daily lives and share social updates. This convenient image acquisition method has led to an exponential increase in multimedia data on users' personal devices, with massive amounts of digital photos covering users' life trajectories at different times, in different geographical locations, and in different scenarios.
[0003] However, in existing image management and display technologies, photo album applications typically present photos using time-based linear lists or file-store-order-based grid tiling. This display method is limited to two-dimensional visual interaction, resulting in a relatively simple format and a lack of spatial depth. How to manage and display these image resources efficiently, intuitively, and engagingly has become a key aspect of improving user experience. Summary of the Invention
[0004] The purpose of this application is to provide an image display method and apparatus that can solve the problem of not being able to manage and display image resources efficiently, intuitively, and in an interesting way.
[0005] In a first aspect, embodiments of this application provide an image display method, including: Receive the first input; In response to the first input, multimodal features of multiple images are extracted and metadata of each image is parsed; wherein, the multiple images are all or some of the images in the first application; The camera pose corresponding to the capture time of each image is determined based on multimodal features and metadata; Multiple images are displayed in a virtual 3D space based on the camera pose.
[0006] Secondly, embodiments of this application provide an image display device, including: The first receiving module is used to receive the first input; An extraction module is used to extract multimodal features from multiple images and parse metadata of each image in response to a first input; wherein the multiple images are all or some of the images in a first application; The first determining module is used to determine the camera pose corresponding to each image capture based on multimodal features and metadata; The first display module is used to display multiple images in a virtual three-dimensional space based on the camera pose.
[0007] Thirdly, embodiments of this application provide an electronic device, which includes a processor and a memory. The memory stores programs or instructions that can run on the processor, and when the programs or instructions are executed by the processor, they implement the steps of the image display method provided in embodiments of this application.
[0008] Fourthly, embodiments of this application provide a readable storage medium storing a program or instructions, which, when executed by a processor, implement the steps of the image display method provided in embodiments of this application.
[0009] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the steps of the image display method provided in embodiments of this application.
[0010] Sixthly, embodiments of this application provide a computer program product, which is stored in a storage medium and executed by at least one processor to implement the steps of the image display method provided in embodiments of this application.
[0011] In this embodiment, a first input is received; in response to the first input, multimodal features of multiple images are extracted and metadata of each image is parsed; wherein the multiple images are all or part of the images in a first application; the camera pose corresponding to the capture of each image is determined based on the multimodal features and metadata; and the multiple images are displayed in a virtual three-dimensional space based on the camera pose. Thus, by determining the camera pose corresponding to the capture of each image using the multimodal features and metadata of multiple images in the first application, and then displaying multiple images in a virtual three-dimensional space based on the camera pose corresponding to the capture of each image, multiple images in the first application can be displayed in a three-dimensional manner, enabling efficient, intuitive, and engaging image management and display. Attached Figure Description
[0012] Figure 1 This is a flowchart illustrating an image display method provided in some embodiments of this application; Figure 2 These are schematic diagrams showing the results of displaying images provided in some embodiments of this application; Figure 3 This is a schematic diagram of the topology of the displayed images provided in some embodiments of this application; Figure 4 These are schematic diagrams of the structure of an image display device provided in some embodiments of this application; Figure 5 These are schematic diagrams of the structure of electronic devices provided in some embodiments of this application; Figure 6These are schematic diagrams of the hardware structure of electronic devices provided in some embodiments of this application. Detailed Implementation
[0013] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0014] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0015] The image display method and apparatus provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0016] It should be noted that the image display method provided in this application can be executed by electronic devices such as mobile phones, tablets, laptops, PDAs, and in-vehicle electronic devices. Some embodiments of this application use electronic devices as the executing entity to illustrate the image display method provided in this application.
[0017] Figure 1 This is a schematic flowchart illustrating an image display method provided in some embodiments of this application. The image display method includes: Step 101: Receive the first input; In some embodiments of this application, the first input is used to trigger the display of multiple images from a first application in a virtual three-dimensional space. The first input in these embodiments includes, but is not limited to, touch input by a user using a touch device such as a finger or stylus to a control used to trigger the display of multiple images from the first application in a virtual three-dimensional space; wherein, touch input includes, but is not limited to, click input, swipe input, etc. Click input can be a single click, double click, or any number of clicks, and can also be a long press or a short press. The first input in these embodiments can be set and modified adaptively according to actual needs.
[0018] In some embodiments of this application, the first application may be an application that includes multiple images, such as a photo album application.
[0019] Step 102: In response to the first input, extract multimodal features from multiple images and parse the metadata of each image; wherein, the multiple images are all or some of the images in the first application; In some embodiments of this application, when parsing the metadata of each image, the Exchangeable Image File Format (EXIT) information of the image file can be read to obtain the image's metadata. The EXIT information includes, but is not limited to: image creation date and time (DateTimeOriginal), camera lens focal length, Global Positioning System (GPS) coordinates of the shooting location, latitude of the shooting location, longitude of the shooting location, camera orientation, camera manufacturer or model (Make / Model), and camera three-dimensional attitude angle.
[0020] In some embodiments of this application, when the GPS coordinates of an image and the image creation date and time are obtained, the GPS coordinates of multiple images can be uniformly converted to a geodetic coordinate system (World Geodetic System 1984, WGS84) or a local tangent plane coordinate system (East-North-Up, ENU), and the image creation date and time can be uniformly converted to Coordinated Universal Time (UTC).
[0021] In some embodiments of this application, for images where the FocalLength is missing in the EXIT information, the camera's field of view can be estimated by the image size.
[0022] In some embodiments of this application, a pre-trained multimodal semantic feature extraction model can be used to convert an image into a high-level embedding vector to obtain the multimodal semantic features of the image; combined with reverse geocoding and a visual scene classification network, the geographic semantic features of the image are generated; image labels and confidence indices are generated through text tag library matching; and a monocular depth estimation neural network is used to generate the relative depth map and normal map of the image. The multimodal semantic feature extraction model can be based on a contrastive language-image pretraining (CLIP) model or a sigmoid loss-based language-image pretraining (SigLIP) model. Reverse geocoding is the process of converting latitude and longitude coordinates into human-readable address or location information (such as streets, cities, and attractions). Combined with Points of Interest (POI) data from the open-source map database (OpenStreetMap, OSM), high-precision geographic information services are achieved. Monocular depth estimation neural networks include, but are not limited to: multi-scale deep stereo (MiDaS) models, dense prediction transformer (DPT) models, absolute depth estimation (ZoeDepth) models, or monocular depth estimation models (Depth Anything) optimized for mobile and edge devices.
[0023] In some embodiments of this application, relative depth maps can be aligned to a uniform metric scale based on camera focal length and ground / horizon detection, ensuring geometric consistency of different images in virtual 3D space.
[0024] In some embodiments of this application, when multiple images are part of a first application, the multiple images can be images selected by the user from images in the first application.
[0025] In some embodiments of this application, before step 102, the image display method provided in this application may further include: deduplicating multiple images.
[0026] When deduplicating multiple images, the similarity between two images can be determined by the hash value corresponding to each image. When the similarity between the images is greater than a first threshold, one of the two images is retained.
[0027] Step 103: Determine the camera pose for each image based on multimodal features and metadata; In some embodiments of this application, when determining the camera pose corresponding to each image based on multimodal features and metadata, the camera pose can first be coarsely located using the absolute position provided by GPS and the gravity direction provided by the Inertial Measurement Unit (IMU). Then, Structure from Motion (SFM) technology is used to extract local image features, perform feature matching and geometric verification, and refine the camera pose calculation. A factor map is then constructed, using visual reprojection error as a visual factor and GPS / IMU readings as prior factors or inertial factors. The camera pose is jointly optimized using a nonlinear least squares method. Scale-invariant feature transform (SIFT) or a deep learning-based local feature detection and description algorithm (SuperPoint) can be used to extract local image features.
[0028] Step 104: Display multiple images in a virtual 3D space based on the camera pose.
[0029] In some embodiments of this application, the virtual three-dimensional space in these embodiments is a three-dimensional space created on the screen of an electronic device using computer technology.
[0030] In some embodiments of this application, in step 104, the image can be mapped as a box in a virtual three-dimensional space, with a photo texture applied to the front and metadata displayed on the back; physically based rendering (PBR) is used to render the image, and Fresnel effect is used to highlight the edges. In the pixel shader, depth maps are used to implement parallax mapping or relief mapping, allowing the user to perceive the three-dimensionality within the photo as the viewing angle moves.
[0031] For example, such as Figure 2 As shown, Figure 2 These are schematic diagrams showing the results of displaying images provided in some embodiments of this application. Figure 2The image shows three images taken at location A of trees and two images taken at location B of a house. The three images at location A include one taken directly in front of the tree, one taken to the right of the tree, and one taken to the left of the tree. The two images at location B include one taken directly in front of the house and one taken to the right of the house. These images are displayed in three dimensions within a virtual 3D space.
[0032] In some embodiments of this application, when displaying an image, the image can also be anchored according to the camera pose; avoidance can be performed when adjacent images surround the thin box; when an image is occluded by other images, the transparency of the occluded image can be reduced or only the outline of the occluded image can be displayed, etc.
[0033] In some embodiments of this application, the direction and intensity of illumination in the virtual three-dimensional space can also be adjusted according to DateTimeOriginal; and the image can be subjected to tone mapping, exposure adaptation, and screen space ambient occlusion (SSAO) using a professional color compilation system (ACES).
[0034] In some embodiments of this application, the range of the view frustum for a future period of time can be predicted based on the user's current image viewing trajectory in the virtual three-dimensional space, and the texture and model data of the image corresponding to the view frustum range can be preloaded; based on the intersection of the image viewing trajectory and the view frustum, the preloaded image is decoded and cached in the background.
[0035] For example, a user takes a series of travel photos using a smartphone, including GPS and EXIF data. The electronic device extracts semantic vectors, determines that the photos have a high overlap rate at a certain scenic spot, triggers the training of point-based volumetric rendering technology, and generates a roamable 3D model of that scenic spot. For scattered photos along the way, a "city street" style background is matched. The user can view the 3D model from all angles and also trace back the journey along the street background.
[0036] For example, photos in a user's album lack GPS and IMU information. The electronic device first determines the general environment using a scene recognition model; for example, if the recognized environment is a "courtyard," it uses SFM to recover the relative positions between photos. If visual overlap is insufficient, the system uses semantic analysis to cluster the photos by time and semantics, and displays them as wall hangings in the generated "virtual living room" scene, constructing a virtual memory space.
[0037] For example, in multi-story indoor scenarios such as shopping malls or homes, electronic devices combine Wi-Fi round-trip time (RTT) signals with visual-inertial odometry (VIO) to distinguish floor heights. During rendering, "elevators" or "stairs" are used as channels connecting different height planes. When a user navigates to the stairs, a floor transition animation is triggered, loading photo data and background from another floor to achieve a vertical spatial narrative.
[0038] In this embodiment, a first input is received; in response to the first input, multimodal features of multiple images are extracted and metadata of each image is parsed; wherein the multiple images are images in a first application; the camera pose corresponding to each image is determined based on the multimodal features and metadata; and the multiple images are displayed in a virtual three-dimensional space based on the camera pose. Thus, multiple images in the first application can be displayed in three dimensions.
[0039] In some embodiments of this application, before step 104, the image display method provided in this application may further include: generating an environmental background image corresponding to multiple images; and displaying the environmental background image in a virtual three-dimensional space.
[0040] In some embodiments of this application, the view frustum overlap rate of the camera within the target area can be determined. When the view frustum overlap rate is less than a preset threshold, a style template corresponding to the semantic label can be selected from the style template library based on the semantic label of the image. An environmental background image is generated based on the style template. The tones of multiple images are clustered to obtain the main tones of the multiple images. Based on the main tones, the tones of the environmental background image are adjusted to blend the atmosphere of the environmental background image with that of the multiple images. When the view frustum overlap rate is greater than or equal to the preset threshold, an environmental background image can be generated using a pre-trained background generation model based on Neural Radiation Field (NeRF) or 3D Gaussian Splatting. The background generation model can be initialized based on sparse point clouds, trained using block-wise training, and image encoded using multi-resolution hash coding. The background generation model can also support streaming loading and real-time rendering.
[0041] In some embodiments of this application, the image display method provided in this application may further include: receiving a second input to a first image; wherein the first image is any one of a plurality of images; in response to the second input, displaying a first topological graph corresponding to the first image; wherein the first topological graph includes K images, K is a positive integer, the K images are images corresponding to K nodes in an undirected graph, the undirected graph is an undirected graph constructed based on the plurality of images; the K nodes are nodes corresponding to edges that satisfy a first condition among the edges formed by the first nodes corresponding to the first images.
[0042] The embodiments of this application do not limit the shape of the topology graph. Any available shape can be applied to the embodiments of this application, such as sphere, ring, spiral, etc.
[0043] In some embodiments of this application, the second input is used to trigger the display of a first topological map corresponding to the first image. The second input in these embodiments includes, but is not limited to, touch input to the first image by a user using a touch device such as a finger or stylus. The second input in these embodiments can be set and modified adaptively according to actual needs.
[0044] In some embodiments of this application, the first condition in these embodiments includes: The weight of the edge is greater than or equal to the first threshold; or, After sorting the edges according to their weights from largest to smallest, the sorting index of the edges is less than or equal to K.
[0045] In some embodiments of this application, before displaying the topology map corresponding to the first image, the image display method provided in this application further includes: determining the weight of the first edge based on the semantic similarity between the second image and the first image, the first distance, and the first time difference, wherein the second image is any image other than the first image among a plurality of images, the first distance is the distance between the shooting location of the second image and the shooting location of the first image, the first time difference is the time difference between the shooting time of the second image and the shooting time of the first image, and the first edge is the edge formed by the node corresponding to the second image and the first node corresponding to the first image.
[0046] In some embodiments of this application, an undirected graph can be constructed that includes multiple nodes corresponding to multiple images, wherein the multiple images correspond one-to-one with the multiple nodes.
[0047] In some embodiments of this application, when determining the weight of the first edge based on the semantic similarity, first distance, and first time difference between the second image and the first image, the weight of the edge can be determined by the following formula (1): (1) In formula (1), Let be the weight of the edge formed by the nodes corresponding to images i and j. Let be the semantic similarity between image i and image j. Let be the distance between the locations where images i and j were captured. Let be the time difference between the capture times of image i and image j. , and As weight, and This is the normalization function.
[0048] In some embodiments of this application, , and It can be dynamically adjusted based on the user interaction context; for example, when it detects that the user is quickly swiping the timeline, Automatically enlarges; when a user stares at a photo for an extended period of time... Automatically enlarges to recommend photos with similar content.
[0049] In some embodiments of this application, after receiving the second input of the first image, the system searches the undirected graph for the K edges with the largest weights among the edges formed by the nodes corresponding to the first image, thereby obtaining the image corresponding to the node of the searched edge. Alternatively, the system can search the undirected graph for edges with weights greater than or equal to a first threshold among the edges formed by the nodes corresponding to the first image, thereby obtaining the image corresponding to the node of the searched edge.
[0050] In some embodiments of this application, when displaying the first topology map corresponding to the first image, the images corresponding to the nodes of the searched edges can be arranged in a preset layout in a virtual three-dimensional space with the first image as the center and the radius mapped according to the weight of the edges, so as to obtain the first topology map corresponding to the first image.
[0051] For example, such as Figure 3 As shown, Figure 3 This is a schematic diagram of the topology of images provided in some embodiments of this application. Figure 3 The topology graph of image i is shown. In addition to image i, the topology graph of image i also includes image a, image b, image c and image d. The weight of the edge formed by the nodes corresponding to image i and image a is 0.9, the weight of the edge formed by the nodes corresponding to image i and image b is 0.8, the weight of the edge formed by the nodes corresponding to image i and image c is 0.7, and the weight of the edge formed by the nodes corresponding to image i and image d is 0.85.
[0052] In some embodiments of this application, the image display method provided in this application further includes: receiving a third input to a third image; wherein the third image is any one of K images of a first topological map; and displaying a second topological map corresponding to the third image in response to the third input.
[0053] In some embodiments of this application, the third input is used to trigger the display of a second topology map corresponding to the third image. The third input in these embodiments includes, but is not limited to, touch input to the third image by a user using a touch device such as a finger or stylus. The third input in these embodiments can be set and modified adaptively according to actual needs.
[0054] In some embodiments of this application, the process of displaying the second topology map corresponding to the third image is similar to the process of displaying the first topology map corresponding to the first image. The embodiments of this application will not be described in detail here. For the specific process, please refer to the description in the process of displaying the first topology map corresponding to the first image.
[0055] For example, when the user clicks the above Figure 3 When dealing with image 'a', the process involves searching for the K edges with the highest weights among the edges formed by the nodes corresponding to image 'a' in the undirected graph, thereby obtaining the images corresponding to the nodes of the searched edges. Centered on image 'a', the process maps radii according to the weights of the edges, and then arranges the images corresponding to the nodes of the searched edges in a virtual 3D space according to a preset layout, resulting in a topology graph corresponding to image 'a'.
[0056] For example, when the user clicks the above Figure 3 When dealing with image a, the algorithm searches the undirected graph for edges whose weights are greater than or equal to a first threshold among the edges formed by the nodes corresponding to image a, thereby obtaining the images corresponding to the nodes of the searched edges. Centered on image a, the algorithm maps the radius according to the weights of the edges and arranges the images corresponding to the nodes of the searched edges in a virtual 3D space according to a preset layout, thus obtaining the topology graph corresponding to image a.
[0057] In some embodiments of this application, an electronic device can transmit multiple images to the cloud, where the cloud extracts multimodal features of the multiple images and parses the metadata of each image; and determines the camera pose corresponding to each image based on the multimodal features and metadata.
[0058] In some embodiments of this application, when an electronic device transmits multiple images to the cloud, it can perform Gaussian blurring or pixelation on faces, license plates, and text in the images; convert GPS coordinates into hexagonal grid indexes or superimpose random noise of Laplacian distribution on the original coordinates to ensure that the cloud cannot reverse-engineer the location where the image was taken; and truncate the time stamp precision of DateTimeOriginal from "seconds" to "minutes" or "hours", etc.
[0059] In some embodiments of this application, the electronic device and the cloud can use Transport Layer Security (TLS) version 1.3 for data transmission and implement Certificate Pinning to prevent data attacks; sensitive fields are encrypted again at the application layer; the electronic device can store keys in a system-level security module, such as a Trusted Execution Environment (TEE), a keystore, or a keychain. The keystore is a system-level component that securely stores and manages encryption keys to protect sensitive keys from leakage or misuse; the keychain is a sensitive data security storage service at the operating system level, used to persistently store confidential information such as passwords, encryption keys, and certificates.
[0060] The image display method provided in this application can be executed by an image display device. This application uses an image display device executing the image display method as an example to illustrate the image display device provided in this application.
[0061] Figure 4 This is a schematic diagram of the structure of an image display device provided in some embodiments of this application. The image display device 300 may include: The first receiving module 401 is used to receive the first input; Extraction module 402 is used to extract multimodal features of multiple images and parse metadata of each image in response to the first input; wherein the multiple images are all or part of the images in the first application; The first determining module 403 is used to determine the camera pose corresponding to each image capture based on multimodal features and metadata; The first display module 404 is used to display multiple images in a virtual three-dimensional space according to the camera pose.
[0062] In this embodiment, a first input is received; in response to the first input, multimodal features of multiple images are extracted and metadata of each image is parsed; wherein the multiple images are all or part of the images in a first application; the camera pose corresponding to the capture of each image is determined based on the multimodal features and metadata; and the multiple images are displayed in a virtual three-dimensional space based on the camera pose. Thus, by determining the camera pose corresponding to the capture of each image using the multimodal features and metadata of multiple images in the first application, and then displaying multiple images in a virtual three-dimensional space based on the camera pose corresponding to the capture of each image, multiple images in the first application can be displayed in a three-dimensional manner, enabling efficient, intuitive, and engaging image management and display.
[0063] In some embodiments of this application, the image display device 400 further includes: The second receiving module is used to receive a second input to the first image; wherein the first image is any one of a plurality of images; The second display module is used to respond to the second input and display a first topological graph corresponding to the first image; wherein the first topological graph includes K images, K is a positive integer, the K images are images corresponding to K nodes in an undirected graph, the undirected graph is an undirected graph constructed based on multiple images; the K nodes are nodes corresponding to the edges that satisfy the first condition among the edges formed by the first nodes corresponding to the first images.
[0064] In some embodiments of this application, the first condition includes: The weight of the edge is greater than or equal to the first threshold; or, After sorting the edges according to their weights from largest to smallest, the sorting index of the edges is less than or equal to K.
[0065] In some embodiments of this application, the image display device 400 further includes: The second determining module is used to determine the weight of the first edge based on the semantic similarity between the second image and the first image, the first distance, and the first time difference. The second image is any image other than the first image among multiple images. The first distance is the distance between the shooting location of the second image and the shooting location of the first image. The first time difference is the time difference between the shooting time of the second image and the shooting time of the first image. The first edge is the edge formed by the node corresponding to the second image and the first node corresponding to the first image.
[0066] In some embodiments of this application, the image display device 400 further includes: The third receiving module is used to receive a third input to the third image; wherein the third image is any one of the K images of the first topological map; The third display module is used to display a second topological map corresponding to the third image in response to the third input.
[0067] In some embodiments of this application, the image display device 400 further includes: The generation module is used to generate environmental background images corresponding to multiple images; The fourth display module is used to display the environmental background image in the virtual three-dimensional space.
[0068] The image display device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the device.
[0069] The image display device in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system.
[0070] The image display device provided in this application embodiment can achieve... Figures 1 to 3 The various processes implemented in the image display method embodiment will not be described again here to avoid repetition.
[0071] Optionally, such as Figure 5 As shown, this application embodiment also provides an electronic device 500, including a processor 501 and a memory 502. The memory 502 stores a program or instructions that can run on the processor 501. When the program or instructions are executed by the processor 501, they implement the various steps of the image display method embodiment provided in this application embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.
[0072] Figure 6 These are schematic diagrams of the hardware structure of electronic devices according to some embodiments of this application.
[0073] The electronic device 600 includes, but is not limited to, components such as: radio frequency unit 601, network module 602, audio output unit 603, input unit 604, sensor 605, display unit 606, user input unit 607, interface unit 608, memory 609, and processor 610.
[0074] Those skilled in the art will understand that the electronic device 600 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 610 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 6 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.
[0075] The user input unit 607 is used to receive the first input; The processor 610 is configured to, in response to a first input, extract multimodal features from multiple images and parse metadata for each image; wherein the multiple images are all or some of the images in a first application; and determine the camera pose corresponding to each image at the time of capture based on the multimodal features and metadata. Display unit 606 is used to display multiple images in a virtual three-dimensional space according to the camera pose.
[0076] In this embodiment, a first input is received; in response to the first input, multimodal features of multiple images are extracted and metadata of each image is parsed; wherein the multiple images are all or part of the images in a first application; the camera pose corresponding to the capture of each image is determined based on the multimodal features and metadata; and the multiple images are displayed in a virtual three-dimensional space based on the camera pose. Thus, by determining the camera pose corresponding to the capture of each image using the multimodal features and metadata of multiple images in the first application, and then displaying multiple images in a virtual three-dimensional space based on the camera pose corresponding to the capture of each image, multiple images in the first application can be displayed in a three-dimensional manner, enabling efficient, intuitive, and engaging image management and display.
[0077] In some embodiments of this application, the user input unit 607 is further configured to: receive a second input to a first image; wherein the first image is any one of a plurality of images; Accordingly, the display unit 606 is further configured to: in response to the second input, display a first topological graph corresponding to the first image; wherein the first topological graph includes K images, K is a positive integer, the K images are images corresponding to K nodes in an undirected graph, the undirected graph is an undirected graph constructed based on multiple images; the K nodes are nodes corresponding to the edges that satisfy the first condition among the edges formed by the first nodes corresponding to the first image.
[0078] In some embodiments of this application, the first condition includes: The weight of the edge is greater than or equal to the first threshold; or, After sorting the edges according to their weights from largest to smallest, the sorting index of the edges is less than or equal to K.
[0079] In some embodiments of this application, the processor 610 is also used for: The weight of the first edge is determined based on the semantic similarity, first distance, and first time difference between the second image and the first image. The second image is any image other than the first image among multiple images. The first distance is the distance between the shooting location of the second image and the shooting location of the first image. The first time difference is the time difference between the shooting time of the second image and the shooting time of the first image. The first edge is the edge formed by the node corresponding to the second image and the first node corresponding to the first image.
[0080] In some embodiments of this application, the user input unit 607 is further configured to: receive a third input to a third image; wherein the third image is any one of the K images of the first topological map; Accordingly, the display unit 606 is also configured to: display a second topological map corresponding to the third image in response to the third input.
[0081] In some embodiments of this application, the processor 610 is further configured to: generate an environmental background image corresponding to multiple images; Accordingly, the display unit 606 is also used to display an environmental background image in a virtual three-dimensional space.
[0082] It should be understood that, in this embodiment, the input unit 604 may include a graphics processing unit (GPU) 6041 and a microphone 6042. The GPU 6041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 606 may include a display panel 6061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 607 includes at least one of a touch panel 6071 and other input devices 6072. The touch panel 6071 is also called a touch screen. The touch panel 6071 may include a touch detection device and a touch controller. Other input devices 6072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.
[0083] The memory 609 can be used to store software programs and various data. The memory 609 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 609 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 609 in this embodiment includes, but is not limited to, these and any other suitable types of memory.
[0084] Processor 610 may include one or more processing units; optionally, processor 610 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 610.
[0085] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the image display method provided in this application and achieve the same technical effect. To avoid repetition, they will not be described again here.
[0086] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.
[0087] This application also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the image display method provided in this application and achieve the same technical effect. To avoid repetition, it will not be described again here.
[0088] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.
[0089] This application also provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the image display method embodiment provided in this application, and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0090] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.
[0091] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the image display methods provided in the various embodiments of this application.
[0092] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. An image display method, characterized in that, The method includes: Receive the first input; In response to the first input, multimodal features of multiple images are extracted and metadata of each image is parsed; wherein, the multiple images are all or some of the images in the first application; The camera pose corresponding to the capture of each image is determined based on the multimodal features and the metadata; Based on the camera pose, the multiple images are displayed in a virtual three-dimensional space.
2. The method according to claim 1, characterized in that, The method further includes: Receive a second input to a first image; wherein the first image is any one of the plurality of images; In response to the second input, a first topological graph corresponding to the first image is displayed; wherein the first topological graph includes K images, where K is a positive integer, the K images are images corresponding to K nodes in an undirected graph, the undirected graph is an undirected graph constructed based on the multiple images; the K nodes are nodes corresponding to the edges that satisfy the first condition among the edges formed by the first nodes corresponding to the first image.
3. The method according to claim 2, characterized in that, The first condition includes: The weight of the edge is greater than or equal to the first threshold; or, After sorting the edges according to their weights from largest to smallest, the sorting index of the edges is less than or equal to K.
4. The method according to claim 3, characterized in that, Before displaying the topology map corresponding to the first image, the method further includes: The weight of the first edge is determined based on the semantic similarity, first distance, and first time difference between the second image and the first image. The second image is any image other than the first image among the plurality of images. The first distance is the distance between the shooting location of the second image and the shooting location of the first image. The first time difference is the time difference between the shooting time of the second image and the shooting time of the first image. The first edge is the edge formed by the node corresponding to the second image and the first node corresponding to the first image.
5. The method according to claim 2, characterized in that, The method further includes: Receive a third input for a third image; wherein the third image is any one of the K images of the first topological graph; In response to the third input, a second topological map corresponding to the third image is displayed.
6. An image display device, characterized in that, The device includes: The first receiving module is used to receive the first input; An extraction module is configured to, in response to the first input, extract multimodal features from multiple images and parse metadata for each image; wherein the multiple images are all or some of the images in the first application; The first determining module is used to determine the camera pose corresponding to each image capture based on the multimodal features and the metadata; The first display module is used to display the multiple images in a virtual three-dimensional space according to the camera pose.
7. The apparatus according to claim 6, characterized in that, The device further includes: The second receiving module is used to receive a second input to the first image; wherein the first image is any one of the plurality of images; The second display module is configured to respond to the second input and display a first topological graph corresponding to the first image; wherein the first topological graph includes K images, where K is a positive integer, the K images are images corresponding to K nodes in an undirected graph, the undirected graph is an undirected graph constructed based on the multiple images; the K nodes are nodes corresponding to the edges that satisfy a first condition among the edges formed by the first nodes corresponding to the first images.
8. The apparatus according to claim 7, characterized in that, The first condition includes: The weight of the edge is greater than or equal to the first threshold; or, After sorting the edges according to their weights from largest to smallest, the sorting index of the edges is less than or equal to K.
9. The apparatus according to claim 8, characterized in that, The device further includes: The second determining module is used to determine the weight of the first edge based on the semantic similarity between the second image and the first image, the first distance, and the first time difference. The second image is any image other than the first image among the plurality of images. The first distance is the distance between the shooting location of the second image and the shooting location of the first image. The first time difference is the time difference between the shooting time of the second image and the shooting time of the first image. The first edge is the edge formed by the node corresponding to the second image and the first node corresponding to the first image.
10. The apparatus according to claim 7, characterized in that, The device further includes: The third receiving module is used to receive a third input to the third image; wherein the third image is any one of the K images of the first topological map; The third display module is used to display a second topological map corresponding to the third image in response to the third input.