XR-based cultural relics display method, equipment and device
Through XR equipment, scanning cultural relics and generating display content related to modern society has solved the problem of single display methods of traditional cultural relics, realizing immersive experience and efficient understanding.
Patent Information
- Application Number
- CN202510848049.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-06-24
AI Technical Summary
The display method of traditional cultural relics is single, and it is difficult for users to understand the cultural value of cultural relics. The existing technology is complex and the display method is not vivid enough.
Scan cultural relics through XR devices, analyze cultural relics information and generate display content related to modern society, use multimodal large language models to generate video and text information, and combine virtual repair technology to improve the display effect.
It realizes users' immersive experience in the museum, lowers the threshold for understanding cultural relics, and improves the fun and image of the display.
Smart Images

Figure CN120355874B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of XR display, and specifically to XR-based cultural relics display methods, equipment and devices. Background Art
[0002] Cultural relics hold high cultural value, serving as historical witnesses, cultural heritage, artistic treasures, and educational vehicles. However, with the changing times, many cultural relics, due to their peculiar shapes and significant differences from modern life, have a high threshold for users to understand them.
[0003] Traditionally, museums display cultural relics individually, accompanied by simple images and text descriptions. Visitors can gain knowledge about the artifacts by viewing these descriptions. However, this display method is relatively simple and not intuitive enough. Furthermore, without a grasp of the historical context of the time, it is difficult for visitors to understand the artifact's origin and purpose, and thus to appreciate its cultural value.
[0004] With the development of technology, users can use smartphones to scan, search for keywords, and obtain additional information about cultural relics online. However, this method is not only complex, requiring users to search for relevant content in a vast amount of resources, but also often limited to text, images, and videos, which is still not very vivid. Summary of the Invention
[0005] To address the above issues, this application proposes an XR-based cultural relic display method, including:
[0006] Scanning and identifying the exhibited cultural relics through the XR device worn by the user to obtain cultural relic information of the exhibited cultural relics;
[0007] Determining, based on the cultural relic information, social relevance information between the exhibited cultural relic and the current era;
[0008] If the degree of association corresponding to the social association information is lower than a preset threshold, generating display content for describing the use of the exhibited cultural relics in the historical era based on the cultural relic information;
[0009] If the degree of association corresponding to the social association information is higher than a preset threshold, generating display content for describing the use of the exhibited cultural relics in the current era based on the social association information and the cultural relics information;
[0010] The display content is displayed to the user through the XR device.
[0011] In one example, determining the social association information between the exhibited cultural relics and the current era based on the cultural relic information specifically includes:
[0012] Determining the image information and introduction information contained therein based on the cultural relic information;
[0013] Performing a similarity comparison based on the image information to select a first object that has an appearance similarity with the exhibited artifact that is higher than a preset threshold and belongs to the current era;
[0014] Outputting a second object having the same purpose and belonging to the current era through a multimodal large language model according to the introduction information;
[0015] According to the first object and the second object, social correlation information between the exhibited cultural relics and the current era is obtained.
[0016] In one example, based on the social association information and the cultural relic information, display content is generated to describe the usage of the exhibited cultural relic in the current era, specifically including:
[0017] Generating video content constraint information based on the social association information; the video content constraint information at least includes: constructing a display video based on the usage scenario of the second object in the current era, and replacing the second object in the display video with the first object and the exhibited cultural relic respectively;
[0018] Corresponding prompt words are generated according to the social association information, the cultural relic information and the video content constraint information, and a display video describing the use of the exhibited cultural relics in the current era is generated through a multimodal large language model.
[0019] In one example, corresponding prompt words are generated based on the social association information, the cultural relic information, and the video content constraint information, and a display video describing the current use of the exhibited cultural relics is generated using a multimodal large language model, specifically including:
[0020] Generate a first text prompt word based on the social association information and the cultural relic information, and output corresponding first text information through a multimodal large language model; the first text information is used to describe the purpose of the cultural relic information in the current era;
[0021] Generating a second text prompt word based on the first text information, and outputting the corresponding second text information through a multimodal large language model; the second text information includes corresponding video creation information when the purpose is converted from text mode to video mode;
[0022] A third text prompt word is generated based on the second text information, the image information, and the video content constraint information, and a corresponding display video is output through a multimodal large language model as display content describing the use of the exhibited cultural relics in the current era.
[0023] In one example, based on the cultural relic information, display content is generated to describe the use of the exhibited cultural relic in the historical era, specifically including:
[0024] Determining the historical information corresponding to the cultural relic in the historical era based on the cultural relic information;
[0025] generating, based on the historical information, display content for describing the use of the exhibited cultural relics in the historical era;
[0026] According to the chronological order of the historical eras, the display contents corresponding to the historical eras are spliced together.
[0027] In one example, the method further includes:
[0028] Determining image information corresponding to the exhibited cultural relics and observation information of the XR device based on the cultural relic information;
[0029] determining a defective area of the exhibited cultural relic according to the image information;
[0030] Performing virtual repair on the exhibited cultural relic according to the damaged area, and generating a virtual repair process corresponding to the observation information;
[0031] The exhibition artifacts are continuously displayed through the first area of the XR device, and the virtual repair process and the exhibition artifacts obtained after the virtual repair process are displayed on the basis of the exhibition artifacts through the second area of the XR device.
[0032] In one example, the method further includes:
[0033] Determine, based on the positioning information and observation information of the XR devices, multiple XR devices currently observing the same exhibited artifact, and group the multiple XR devices into a dynamic device set;
[0034] If it is determined that a specified condition exists according to the current user's XR device, then selecting an XR device of another user that does not have the specified condition from the dynamic device set as a shared device;
[0035] The display content in the shared device is shared with the current user's XR device until the specified situation is eliminated.
[0036] In one example, if a specified condition exists based on the current user's XR device, then in the dynamic device set, XR devices of other users that do not have the specified condition are selected as shared devices, specifically including:
[0037] If it is determined based on the observation information of the current user's XR device that there is an occlusion area on the observation path, and more than a preset proportion of XR devices in the dynamic device set have occlusion areas on their observation paths, then it is considered that the first specified situation exists;
[0038] If the transmission speed between the current user's XR device and the cloud server is lower than the preset speed, it is considered that the second specified situation exists;
[0039] For the first specified situation, select an XR device of another user that is closest to the exhibited cultural relics and does not meet the first specified situation as a shared device;
[0040] For the second specified situation, an XR device of another user that is closest to the current user and does not meet the second specified situation is selected as the shared device.
[0041] On the other hand, this application also proposes an XR-based cultural relic display device, including:
[0042] at least one processor; and,
[0043] a memory communicatively connected to the at least one processor; wherein,
[0044] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the XR-based cultural relic display method as described in any of the above examples.
[0045] On the other hand, this application also proposes an XR-based cultural relic display device, including:
[0046] A scanning module, which scans and identifies exhibited cultural relics through the XR device worn by the user to obtain cultural relic information of the exhibited cultural relics;
[0047] a correlation module, which determines, based on the cultural relic information, social correlation information between the exhibited cultural relic and the current era;
[0048] A first content generation module generates, based on the cultural relic information, display content describing the use of the exhibited cultural relic in the historical era if the degree of association corresponding to the social association information is lower than a preset threshold;
[0049] The display module displays the display content to the user through the XR device.
[0050] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to implement the XR-based cultural relic display method as described in any of the above examples.
[0051] The XR-based cultural relics display method proposed in this application can bring the following beneficial effects:
[0052] Compared to traditional solutions where users search independently through smartphones, intelligent context reconstruction through XR devices eliminates the need for users to tediously scan and search using smartphones, and presentation methods are no longer limited to text, images, or videos. This lowers the barrier to understanding cultural relics while also increasing the fun of their display. When viewers scan cultural relics wearing XR devices, the system analyzes their relevance to modern life and, for low-relevance artifacts, generates dynamic scenes to restore ancient usage scenarios. This allows users to immerse themselves in the museum's interpretation of cultural relics, making the display more vivid and lifelike. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0054] Figure 1 A schematic diagram of the structure of a system on which the embodiments of the present application can be run;
[0055] Figure 2 Schematic diagram of the process of the XR-based cultural relics display method in an embodiment of the present application;
[0056] Figure 3 This is a schematic diagram of a process for determining social association information in one scenario in an embodiment of the present application;
[0057] Figure 4 This is a flowchart of generating display content when the degree of association is low in one scenario in an embodiment of the present application;
[0058] Figure 5 This is a schematic diagram of a process for virtual restoration of cultural relics in one scenario in an embodiment of the present application;
[0059] Figure 6 This is a flowchart showing a special situation in one embodiment of the present application;
[0060] Figure 7 This is a schematic diagram of the module structure of an XR device in one scenario in an embodiment of the present application;
[0061] Figure 8 This is a schematic diagram of an XR-based cultural relic display device in one scenario in an embodiment of the present application;
[0062] Figure 9 This is a schematic diagram of an XR-based cultural relic display device in one scenario in an embodiment of the present application. DETAILED DESCRIPTION
[0063] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0064] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.
[0065] The XR-based cultural relics display method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Figure 1 As shown, the application environment may include: an XR device 101 worn by a user 107 , a communication network 102 , a cloud server 103 , a text resource server 104 , an image resource server 105 , and a database 106 .
[0066] XR stands for Extended Reality (XR), which is a general term for Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR). XR devices can be configured as VR glasses, AR glasses, MR glasses, and other devices. For ease of description, the embodiments herein are explained using XR glasses as an example. The communication network 102 can serve as a data transmission channel, providing a communication link for communication between the XR device 101 and the cloud server 103. This allows the XR device 101 to collect various types of sound signals, viewing angle information, image information, and other information, and transmit them to the cloud server 103 via the communication network 102. The XR device 101 can also feed back information processed by the cloud server 103 to the XR device 101 for rendering and display. Cloud server 103 provides various services. To improve processing efficiency, it connects to text resource server 104 and image resource server 105. After receiving information collected by XR device 101 via communication network 102, cloud server 103 can retrieve and process the text and image information related to cultural relics. After processing the relevant information, cloud server 103 combines the results of the processing of the cultural relic-related information by XR device 101 with the results, and transmits them to XR device 101 via communication network 102 for presentation to the user.
[0067] The database 106 is connected to the text resource server 104 and the image resource server 105, respectively, and can be used to store and manage the text information of the text resource server 104 and the image information of the image resource server 105. The database 106 can be integrated on the cloud server 103 or placed on other network servers.
[0068] like Figure 2 As shown, the embodiment of the present application provides an XR-based cultural relic display method, including:
[0069] S201: Scan and identify the exhibited cultural relics through the XR device worn by the user to obtain cultural relic information of the exhibited cultural relics.
[0070] A cloud server is set up, which can process corresponding text information and image information, and store data such as cultural relics 2D images, 3D models, historical data, material spectrum data, etc. through the database, and can interact with XR glasses to realize corresponding data transmission and processing functions.
[0071] Cultural relic information can include images (including 2D and 3D images, etc.) of exhibited cultural relics (hereinafter referred to as cultural relics for ease of description), nameplate information, and introductory information. Introductory information can be obtained from the text description next to the cultural relic. Furthermore, this introductory information can be supplemented by accessing records stored online or in a cloud-based database, for example, by adding information such as the cultural relic's age, name, and instructions for use.
[0072] At the same time, a knowledge parsing engine can be set up, which can combine with the knowledge graph to mark the names of cultural relics components, craftsmanship techniques, and historical events in real time as cultural relic information.
[0073] Generally speaking, the user's movement speed can be judged based on the displacement speed of the XR device. When the user is basically in a stationary state, it is considered that the user is interested in the cultural relics he is facing. The user is asked through the XR device, triggering the corresponding scanning and recognition process.
[0074] By scanning cultural relics information, nameplate information, introduction information, etc. through XR glasses, users can trigger the cultural relics scanning function of the XR glasses while wearing XR glasses to view the exhibition and obtain relevant information about the cultural relics.
[0075] The user wears XR glasses (for example, with a 53° field of view) with a capture resolution of 1280*800. Using the XR glasses' multi-angle cameras, the user activates the XR glasses' scanning function, gazes at the cultural relic, and walks around it, completing a 360-degree scan to capture data on the artifact's surface texture, patterns, and designs. The user then aims the camera at the nameplate to capture the plaque's information. Of course, if there are too many people or the venue is limited, it's possible to scan only a portion of the image or only in-situ, obtaining the corresponding two-dimensional image.
[0076] After sourcing the entire cultural relic information retrieval library, preliminary cultural relic information retrieval and matching are completed, and digital cultural relic objects of all matching results are generated in the digital space to facilitate 360-degree surround display of search results.
[0077] S202: Determine the social association information between the exhibited cultural relics and the current era based on the cultural relic information.
[0078] The current era can be defined based on human needs. For example, the current era can be defined as the time after a certain time point (for example, a specific year), or the current era can be defined as the time within a certain period of time (for example, set to 10 years) from the current time point.
[0079] Social relevance information refers to the relationship between the exhibited artifact and objects still in daily use today. This relevance information can be about appearance, usage, etc.
[0080] Specifically, if Figure 3 As shown in FIG, based on the cultural relic information, the image information and introduction information contained therein are determined, wherein the image information is in image mode, and the introduction information is in text mode.
[0081] A similarity comparison is performed based on the image information to select the first object from the current era whose appearance similarity to the exhibited artifact exceeds a preset threshold. The comparison library can be a pre-set library of images from everyday items or an open-source image library available online. Since comparisons are typically large-scale, a perceptual hashing-based comparison method can be used to compress images into fixed-length hash values, allowing for rapid similarity comparisons using Hamming distance.
[0082] Based on the introductory information, a multimodal large language model outputs a second object with the same purpose and belonging to the current era. For example, the prompt might be, "You are a senior historical archaeologist who is also well-versed in contemporary life. I am given an artifact with the introductory information {information}. Please find objects still used today that have the same purpose and quantify the similarity of their uses, with 100 being the highest score." In this case, objects with scores above a certain point can be used as the second object. The content within the {} symbols is a variable and is added based on the actual situation.
[0083] At this point, information about the social relevance of the exhibited artifacts to the current era is obtained based on the first and second objects. The first object's appearance is similar to that of the artifact, while the second object's use is similar to that of the artifact. Therefore, the artifact can be described directly based on these two aspects (there can be multiple first and second objects, but for ease of calculation and user description, only one of each can be selected). These first and second objects can then be displayed on the XR glasses to help the user quickly understand the artifact.
[0084] S203: If the association degree corresponding to the social association information is lower than a preset threshold, generating display content for describing the use of the exhibited cultural relics in the historical era based on the cultural relic information.
[0085] The degree of relevance between the cultural relic and modern society is determined based on the social relevance information corresponding to the cultural relic information. The degree of relevance can be obtained by weighted summation of the similarity scores and quantitative scores corresponding to the first and second objects.
[0086] The display content may include text, images, videos, etc., and its main purpose is to explain the relevant information of the cultural relic to users.
[0087] If the degree of relevance falls below a preset threshold, it indicates that users are having difficulty quickly understanding the artifact. In this case, content is generated based on the artifact's information to illustrate its historical uses. For example, corresponding prompts (including key words, scene descriptions, character settings, and action requirements) are generated based on the artifact's information. These prompts are then fed into a multimodal large language model, which then outputs and displays images or videos of the artifact's historical usage, helping users understand its purpose in society at the time.
[0088] For example, a user wearing XR glasses visits an art museum and encounters a landscape painting. The XR glasses' camera scans the painting, then the nameplate next to it, identifying the brushstroke style of the painting as the artist's preferred style. Based on the painting's name on the nameplate, a video is generated detailing the artist's inspiration for the painting, the creative process, and the subsequent transfer of the painting to the museum, telling the story of the painting's creation and transmission to the present day. At this point, the use of the exhibited artifact is no longer limited to its appreciation as a painting in itself, but can also encompass the story behind its creation.
[0089] For example, when introducing the history of artifacts, after identifying the shape, texture, and color of the artifact, combined with the information on the museum nameplate, we can obtain information such as the age, name, and usage of the artifact, and use the large model capability to create a video to explain the corresponding artifact.
[0090] Furthermore, if Figure 4 As shown in the figure, based on the information of the cultural relic, the historical information corresponding to the historical era is determined. For example, through multimodal large language models, networks, introduction information, etc., the cultural relic information at each historical time point (including its use and the events it experienced at each historical time point) is obtained.
[0091] Based on historical information, we generate display content that describes the historical uses of exhibited artifacts. For each historical era, we generate prompt words based on the artifact information corresponding to that era. These are then fed into a multimodal large language model to generate a video displaying the artifacts at that historical era.
[0092] According to the chronological order of the historical eras, the display content corresponding to the historical eras is spliced. After obtaining the display videos of each historical era, the display videos of each node are spliced in chronological order to obtain the complete display video and play it.
[0093] In another case, if the degree of association corresponding to the social association information is higher than a preset threshold, it is believed that with the cooperation of appropriate display content, the user's understanding of the cultural relic can be accelerated. Therefore, at this time, based on the social association information and cultural relic information, display content is generated to describe the use of the exhibited cultural relics in the current era.
[0094] Specifically, for cultural relics with a high degree of correlation corresponding to social correlation information, if the above-mentioned method is still used to explain them, it may not be vivid enough, and users still find it difficult to associate the cultural relics with the society of the current era. Therefore, the use of the cultural relics in the current era can be used to generate display content for display, making the entire demonstration process more vivid.
[0095] Based on the social association information, video content constraint information is generated. The video content constraint information refers to the constraints required when generating the corresponding display video. The constraints can be added to the corresponding prompt words to generate the display video.
[0096] The video content constraint information at least includes: constructing a display video according to a usage scenario of the second object in the current era, and replacing the second object in the display video with the first object and the exhibition artifacts respectively.
[0097] Because the second object and the artifact have similar uses, the video is first presented based on their construction. For example, if the artifact is a three-legged jue cup, primarily used for drinking, the second object could be a contemporary wine glass, and a drinking scene could be constructed based on the wine glass. The first object could be a three-legged mug, a modern artifact.
[0098] When the first object and the exhibition cultural relic are replaced respectively, two sub-videos can be generated respectively. For example, in the first half of the constructed display video, the user uses the first object to replace the second object, and in the second half of the display video, the user uses the exhibition cultural relic to replace the second object. The first half and the second half are then spliced together to obtain the display video. In this way, in the display video, not only the use of cultural relics is vividly displayed to the user through modern scenes, but also ancient and modern objects with similar appearances can appear in the same frame, increasing the user's sense of immersion in the scene.
[0099] At this time, corresponding prompt words are generated according to the social association information, cultural relic information and the video content constraint information, and a display video describing the use of the exhibited cultural relics in the current era is generated through a multimodal large language model.
[0100] Furthermore, when generating prompts, a first text prompt can be generated based on the social relevance information and the cultural relic information. The corresponding first text information is then outputted through a multimodal large language model. The first text information is used to describe the current use of the cultural relic information. For example, the first text prompt could be, "You are a senior historical cultural relic expert who is also well-versed in contemporary life. You are given a cultural relic. Its introduction information is {information}, and its social relevance to current society includes {relevance}. You are responsible for associating this social relevance with the current society and in which scenarios it can serve similar purposes."
[0101] A second text prompt is generated based on the first text information, and the corresponding second text information is output through a multimodal large language model. The second text information includes video creation information corresponding to the conversion of the use from text mode to video mode. For example, the second text prompt could be "You are a senior historical cultural relic expert who is also well-versed in contemporary life knowledge and an expert in photography and video creation. You are given a text message {use} describing the use of cultural relics in the current era. You are required to generate a video creation information for the video. This video creation information includes but is not limited to: main plot structure, storyboard design, visual style, sound requirements, and technical constraints."
[0102] A third text prompt is generated based on the second text and image information, and a corresponding display video is output through a multimodal large language model as display content describing the current use of the exhibited cultural relics. For example, the third text prompt could be "You are a senior video creator. Here are the relevant requirements for video creation information {create}, here is the image information of the cultural relic in the video creation {image}, and here is the video content constraint information {constraint}. Based on this video creation information and image information, you will generate a video to demonstrate the current use of this cultural relic in society, while also meeting the requirements of the video content constraint information."
[0103] S204: Display the display content to the user through the XR device.
[0104] If a user has questions about the artifact during the display, they can voice-inquire about it. This question, along with the artifact information, is then fed into the multimodal large language model. For example, regarding the landscape painting mentioned above, a user might ask, "What creative techniques were used in this work?" The large model, upon receiving the question, combines the relevant introductory information identified by the XR scan with the generated video to explain the display to the user. At this point, the message "Video loading..." is displayed.
[0105] During the loading process of the display video, the display video is generated by inputting the large model through the shot content design. After the generation is completed, the loading state is converted to the video screen image, and after displaying "Start automatic play", the video display begins.
[0106] Through intelligent context reconstruction, the threshold for understanding cultural relics is lowered. When visitors wear XR devices to scan cultural relics, the system analyzes their relevance to modern life. For highly relevant cultural relics, real-time comparison data with modern products is overlaid. For less relevant cultural relics, dynamic scenes are generated to restore the ancient usage context. This allows users to immerse themselves in the interpretation of cultural relics while browsing in the museum and increases their interest.
[0107] In one embodiment, virtual restoration can also be performed through XR to show users the complete appearance of the cultural relics.
[0108] Specifically, if Figure 5 As shown, the image information corresponding to the exhibited artifacts and the observation information of the XR device are determined based on the artifact information. Using the XR device or pre-installed laser equipment, laser scanning and multispectral imaging can be used to construct a high-precision 3D model of the artifact, storing material, texture, and historical context information in the cloud. Observation information can include the viewing angle and distance of the exhibited artifact.
[0109] Determine the defective areas of exhibited artifacts based on image information. For example, a lightweight model can be deployed on XR glasses to detect defective areas in real time. The lightweight model backbone network can be MobileNetV3-Small, with a corresponding detection head based on YOLOv5s (including 3D convolutional layers, depthwise separable convolutional layers, and a lightweight attention mechanism) to detect defective areas and output their coordinates and type.
[0110] Virtually repair exhibited artifacts based on damaged areas, and generate a virtual repair process corresponding to the observed information. For example, Generative Adversarial Networks (GANs) can be used to match historical data to generate a virtual 3D model of the repaired artifact. Alternatively, a multimodal large language model can be used to input exhibited artifacts marked with damaged areas and issue instructions to generate a virtual repair process based on the current observation information (including observation angle and distance). In this case, the prompt could be, "You are a senior expert in artifact restoration. The image information of the current artifact is {image}, which has the damaged area marked. Using your expertise, generate a video corresponding to the repair process of this artifact based on the current observation information {observe}."
[0111] Virtual repair can be used for damaged objects or paintings. After identifying the damaged area, the virtual repair process is generated by filling it in. It can also be used for faded murals. After identifying the faded area, the original color simulation effect is superimposed as a virtual repair process.
[0112] At this time, the exhibition relics are continuously displayed through the first area of the XR device, and the virtual repair process and the exhibition relics obtained after the virtual repair process are displayed based on the exhibition relics through the second area of the XR device.
[0113] Among them, after the virtual repair process is generated, the virtual repair model can be aligned with the real cultural relic space through SLAM technology, which can support the seamless visual integration of virtual cultural relics and real cultural relics, and also support the display of virtual cultural relics in space in accordance with real cultural relics.
[0114] When the XR device is an XR pair of glasses, one lens can be used as the first area and the other lens as the second area. This allows users to simultaneously display the artifact's current appearance and its restoration process, increasing their interest in the restoration process. Furthermore, because the restoration process is generated based on current observations, the content seen in both lenses is essentially the same, without affecting the user's viewing experience.
[0115] In one embodiment, when visiting places such as museums, there may be some special circumstances that make it difficult for the XR device to display the content normally.
[0116] Based on this, Figure 6 As shown, based on the positioning and observation information of the XR devices, multiple XR devices currently observing the same exhibited artifact are identified and grouped into a dynamic device set. Based on the positioning information, XR devices in the same area are identified, and then based on the observation information, the XR devices observing the same artifact are identified.
[0117] The XR devices in the dynamic device set are dynamic. Each exhibition artifact corresponds to a dynamic device set. When the user goes to see other artifacts, he or she will automatically leave the dynamic device set corresponding to the current exhibition artifact.
[0118] At this point, if the specified condition is detected based on the current user's XR device, the XR device of another user that does not have the specified condition is selected from the dynamic device set as the shared device. The displayed content on the shared device is shared with the current user's XR device until the specified condition is resolved.
[0119] By sharing the display content with the current user through a shared device, the current user's XR device can quickly avoid the specified situation when encountering it, without affecting the current user's observation experience.
[0120] Furthermore, with respect to the observation situation, if it is determined based on the observation information of the current user's XR device that there is an occlusion area on the observation path, and there is an occlusion area on the observation path of XR devices exceeding a preset proportion in the dynamic device concentration, then it is considered that the first specified situation exists.
[0121] If only the current user has the problem of blocked area, the user can be reminded to move slightly. However, if multiple users have the problem of blocked area, it is considered that there are too many people visiting, resulting in too large a proportion of blocked area.
[0122] At the same time, if the transmission speed between the current user's XR device and the cloud server is lower than the preset speed, it is considered that the second specified situation exists. At this time, it may be caused by an abnormality in the current environment or the transmission module of the device.
[0123] In this case, for the first specified scenario, the XR device closest to the exhibited artifact and not in the first specified scenario is selected as the shared device. This shared device is very close to the exhibited artifact and can fully observe the shape of the exhibited artifact, so it is selected as the shared device for sharing the displayed content.
[0124] For the second specified scenario, the XR device closest to the current user, and not associated with another user in the second specified scenario, is selected as the shared device. For XR devices with slower transmission speeds, the closest XR device is selected as the shared device. D2D communication is then established between the shared device and the current XR device (for example, via Wi-Fi Direct / Ultra Wideband). This ensures that the current XR device can properly play the displayed content, improving the user experience.
[0125] The present application also provides an XR device, such as Figure 7 As shown, the XR device can be XR glasses, including a processor 701, a memory 702, a shooting module 703, a display module 704, an interaction module 705 and a wireless communication module 706.
[0126] The memory 702 can be used to store computer programs, such as software programs and modules of application software. The processor 701 executes various functional applications and data processing by running the computer programs stored in the memory 702, such as implementing the method provided in the embodiment of the present application.
[0127] The shooting module 703 is mainly responsible for shooting the introduction information and image information of the cultural relics. The shooting module can be a camera.
[0128] The display module 704 is used to display the cultural relic restoration results obtained by the method provided in the embodiment of the present application.
[0129] The interaction module 705 includes a microphone array, etc., which is used to realize the interaction between the XR glasses and the user, making it convenient for the user to input commands.
[0130] The wireless communication module 706 includes: 4G / 5G full network access, WiFi, Bluetooth, etc.
[0131] It can be understood by those skilled in the art that Figure 7 The structure shown is only for illustration and does not limit the structure of XR glasses. Figure 7 More or fewer components than shown, or with Figure 7 Different configurations shown.
[0132] The following is a brief introduction to the application scenarios of XR glasses.
[0133] During the use of XR glasses, when the user wears XR glasses to browse the museum, based on the user's trigger or pre-setting, the XR glasses scan the cultural relic information of the exhibited cultural relics, including introduction information, image information, etc., and transmit the cultural relic information to the cloud server. The cloud server determines the social relevance information of the exhibited cultural relics with the current era, and based on the level of social relevance information, generates different display content, and feeds it back to the XR glasses. The XR glasses display the display content to the user so that the user can understand the exhibited cultural relics more vividly.
[0134] like Figure 8 As shown, the embodiment of the present application also provides an XR-based cultural relic display device, including:
[0135] The scanning module 801 scans and identifies the exhibited cultural relics through the XR device worn by the user to obtain cultural relic information of the exhibited cultural relics;
[0136] The association module 802 determines the social association information between the exhibited cultural relics and the current era based on the cultural relic information;
[0137] The first content generation module 803 generates, based on the cultural relic information, display content describing the use of the exhibited cultural relic in the historical era if the degree of association corresponding to the social association information is lower than a preset threshold;
[0138] The display module 804 displays the display content to the user through the XR device.
[0139] Optionally, the association module 802 determines the image information and introduction information contained therein based on the cultural relic information;
[0140] Performing a similarity comparison based on the image information to select a first object that has an appearance similarity with the exhibited artifact that is higher than a preset threshold and belongs to the current era;
[0141] Outputting a second object having the same purpose and belonging to the current era through a multimodal large language model according to the introduction information;
[0142] According to the first object and the second object, social correlation information between the exhibited cultural relics and the current era is obtained.
[0143] Optionally, the system further includes: a second content generation module 805, which generates video content constraint information based on the social association information if the association degree corresponding to the social association information is higher than a preset threshold; the video content constraint information at least includes: constructing a display video based on the usage scenario of the second object in the current era, and replacing the second object in the display video with the first object and the exhibition artifact respectively;
[0144] Corresponding prompt words are generated according to the social association information, the cultural relic information and the video content constraint information, and a display video describing the use of the exhibited cultural relics in the current era is generated through a multimodal large language model.
[0145] Optionally, the second content generation module 805 generates a first text prompt word based on the social association information and the cultural relic information, and outputs corresponding first text information through a multimodal large language model; the first text information is used to describe the purpose of the cultural relic information in the current era;
[0146] Generating a second text prompt word based on the first text information, and outputting the corresponding second text information through a multimodal large language model; the second text information includes corresponding video creation information when the purpose is converted from text mode to video mode;
[0147] A third text prompt word is generated based on the second text information, the image information, and the video content constraint information, and a corresponding display video is output through a multimodal large language model as display content describing the use of the exhibited cultural relics in the current era.
[0148] Optionally, the first content generation module 803 determines the historical information corresponding to the cultural relic in the historical era according to the cultural relic information;
[0149] generating, based on the historical information, display content for describing the use of the exhibited cultural relics in the historical era;
[0150] According to the chronological order of the historical eras, the display contents corresponding to the historical eras are spliced together.
[0151] Optionally, the method further includes: a virtual repair module 806 for determining image information corresponding to the exhibited cultural relic and observation information of the XR device according to the cultural relic information;
[0152] determining a defective area of the exhibited cultural relic according to the image information;
[0153] Performing virtual repair on the exhibited cultural relic according to the damaged area, and generating a virtual repair process corresponding to the observation information;
[0154] The exhibition artifacts are continuously displayed through the first area of the XR device, and the virtual repair process and the exhibition artifacts obtained after the virtual repair process are displayed on the basis of the exhibition artifacts through the second area of the XR device.
[0155] Optionally, the system further includes: a device sharing module 807, which determines, based on the positioning information and observation information of the XR device, multiple XR devices currently observing the same exhibited artifact, and groups the multiple XR devices into a dynamic device set;
[0156] If it is determined that a specified condition exists according to the current user's XR device, then selecting an XR device of another user that does not have the specified condition from the dynamic device set as a shared device;
[0157] The display content in the shared device is shared with the current user's XR device until the specified situation is eliminated.
[0158] Optionally, the device sharing module 807 determines that a first specified condition exists if, based on observation information of the current user's XR device, it is determined that an occlusion area exists on the observation path, and if more than a preset proportion of XR devices in the dynamic device set have occlusion areas on their observation paths;
[0159] If the transmission speed between the current user's XR device and the cloud server is lower than the preset speed, it is considered that the second specified situation exists;
[0160] For the first specified situation, select an XR device of another user that is closest to the exhibited cultural relics and does not meet the first specified situation as a shared device;
[0161] For the second specified situation, an XR device of another user that is closest to the current user and does not meet the second specified situation is selected as the shared device.
[0162] like Figure 9 As shown, the embodiment of the present application further provides an XR-based cultural relic display device, including:
[0163] at least one processor; and,
[0164] a memory communicatively connected to the at least one processor; wherein,
[0165] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the XR-based cultural relic display method as described in any of the above embodiments.
[0166] An embodiment of the present application further provides a non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to implement the XR-based cultural relic display method as described in any of the above embodiments.
[0167] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the device and medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.
[0168] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects to their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.
[0169] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A method for displaying cultural relics based on XR, characterized in that: include: Scanning and identifying the exhibited cultural relics through the XR device worn by the user to obtain cultural relic information of the exhibited cultural relics; Determining the social relevance of the exhibited cultural relics to the current era based on the cultural relic information, specifically including: determining the image information and introduction information contained therein based on the cultural relic information; Performing a similarity comparison based on the image information to select a first object that has an appearance similarity with the exhibited artifact that exceeds a preset threshold and belongs to the current era; outputting a second object that has the same purpose and belongs to the current era through a multimodal large language model based on the introduction information; and obtaining social association information between the exhibited artifact and the current era based on the first and second objects; If the degree of association corresponding to the social association information is lower than a preset threshold, then generating display content describing the use of the exhibited cultural relics in the historical era based on the cultural relic information; the degree of association is obtained by weighted summation of the similarity scores and quantitative scores corresponding to the first object and the second object; The display content is displayed to the user through the XR device.
2. The XR-based cultural relics display method according to claim 1, characterized in that: After determining the social association information between the exhibited cultural relics and the current era based on the cultural relic information, the method further includes: If the degree of association corresponding to the social association information is higher than a preset threshold, generating video content constraint information based on the social association information; the video content constraint information at least includes: constructing a display video based on the usage scenario of the second object in the current era, and replacing the second object in the display video with the first object and the exhibited cultural relic respectively; Corresponding prompt words are generated according to the social association information, the cultural relic information and the video content constraint information, and a display video describing the use of the exhibited cultural relics in the current era is generated through a multimodal large language model.
3. The XR-based cultural relics display method according to claim 2, characterized in that: Generate corresponding prompt words based on the social association information, the cultural relic information, and the video content constraint information, and generate a display video describing the current use of the exhibited cultural relics through a multimodal large language model, specifically including: Generate a first text prompt word based on the social association information and the cultural relic information, and output corresponding first text information through a multimodal large language model; the first text information is used to describe the purpose of the exhibited cultural relic in the current era; Generating a second text prompt word based on the first text information, and outputting the corresponding second text information through a multimodal large language model; the second text information includes video creation information corresponding to converting the use of the exhibited cultural relics in the current era from text mode to video mode; A third text prompt word is generated based on the second text information, the image information, and the video content constraint information, and a corresponding display video is output through a multimodal large language model as display content describing the use of the exhibited cultural relics in the current era.
4. The XR-based cultural relics display method according to claim 1, characterized in that: Based on the cultural relic information, display content is generated to describe the use of the exhibited cultural relic in the historical era, specifically including: Determining the historical information corresponding to the cultural relic in the historical era based on the cultural relic information; generating, based on the historical information, display content for describing the use of the exhibited cultural relics in the historical era; According to the chronological order of the historical eras, the display contents corresponding to the historical eras are spliced together.
5. The XR-based cultural relics display method according to claim 1, characterized in that: The method further comprises: Determining image information corresponding to the exhibited cultural relics and observation information of the XR device based on the cultural relic information; determining a defective area of the exhibited cultural relic according to the image information; Performing virtual repair on the exhibited cultural relic according to the defective area, and generating a virtual repair process corresponding to the observation information; The exhibition artifacts are continuously displayed through the first area of the XR device, and the virtual repair process and the exhibition artifacts obtained after the virtual repair process are displayed on the basis of the exhibition artifacts through the second area of the XR device.
6. The XR-based cultural relics display method according to claim 1, characterized in that: The method further comprises: Determine, based on the positioning information and observation information of the XR devices, multiple XR devices currently observing the same exhibited artifact, and group the multiple XR devices into a dynamic device set; If it is determined that a specified condition exists according to the current user's XR device, then selecting an XR device of another user that does not have the specified condition from the dynamic device set as a shared device; The display content in the shared device is shared with the current user's XR device until the specified situation is eliminated.
7. The XR-based cultural relics display method according to claim 6, characterized in that: If a specified condition exists according to the current user's XR device, then selecting an XR device of another user that does not have the specified condition from the dynamic device set as a shared device specifically includes: If it is determined based on the observation information of the current user's XR device that there is an occlusion area on the observation path, and more than a preset proportion of XR devices in the dynamic device set have occlusion areas on their observation paths, then it is considered that the first specified situation exists; If the transmission speed between the current user's XR device and the cloud server is lower than the preset speed, it is considered that the second specified situation exists; For the first specified situation, select an XR device of another user that is closest to the exhibited cultural relics and does not meet the first specified situation as a shared device; For the second specified situation, an XR device of another user that is closest to the current user and does not meet the second specified situation is selected as the shared device.
8. An XR-based cultural relic display device, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the XR-based cultural relic display method as described in any one of claims 1 to 7.
9. An XR-based cultural relic display device, characterized in that: include: A scanning module, which scans and identifies exhibited cultural relics through the XR device worn by the user to obtain cultural relic information of the exhibited cultural relics; The association module determines, based on the cultural relic information, social association information between the exhibited cultural relic and the current era, specifically comprising: determining, based on the cultural relic information, image information and introduction information contained therein; performing a similarity comparison based on the image information to select a first object having an appearance similarity with the exhibited cultural relic exceeding a preset threshold and belonging to the current era; outputting, based on the introduction information, a second object having the same purpose and belonging to the current era through a multimodal large language model; and obtaining, based on the first and second objects, social association information between the exhibited cultural relic and the current era. A first content generation module generates, based on the cultural relic information, display content describing the historical uses of the exhibited cultural relic if the degree of association corresponding to the social association information is lower than a preset threshold; the degree of association being obtained by weighted summation of the similarity scores and quantitative scores corresponding to the first and second objects; The display module displays the display content to the user through the XR device.
Citation Information
Patent Citations
Multi-modal recognition algorithm based on images and characters
CN119380346A
Method for constructing cultural relic knowledge organization, expression and characterization model for Song charm themes
CN120146164A