Method, equipment and device for displaying cultural relics based on XR
Through XR equipment scanning and identifying cultural relics and generating dynamic display content, the complex and single traditional display methods are solved, and the user immersive explanation of cultural relics and the interest of improving users.
Patent Information
- Application Number
- CN202510848049.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-24
AI Technical Summary
The traditional cultural relics display methods are complex and the display methods are single, making it difficult for users to intuitively understand the cultural value and uses of cultural relics.
The exhibition cultural relics are scanned and identified through the XR device worn by the user, and cultural relics information is obtained, and dynamic display content is generated based on the information, including the use display of the current era and historical era. The multimodal large language model is used to generate video and text prompt words, and combined with intelligent situational reconstruction technology, it provides immersive explanations.
It lowers the threshold for understanding cultural relics, improves the fun and image of the display process, and allows users to more vividly appreciate the cultural value and historical background of cultural relics.
Smart Images

Figure CN120355874A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of XR display, and particularly to a method, device, and apparatus for displaying cultural relics based on XR. Background Art
[0002] Cultural relics can serve as historical witnesses, cultural heritages, artistic treasures, and educational carriers, etc., and have high cultural value. However, with the passage of time, due to their peculiar shapes and large differences from modern life, it is difficult for users to understand many cultural relics.
[0003] In traditional solutions, museums display cultural relics separately and attach simple images and text introductions beside them. Users can obtain corresponding cultural relic knowledge by viewing this introduction information. However, this display method is relatively single and not intuitive enough. Moreover, if users do not understand the historical knowledge of that time, it is difficult to understand the origin, use, etc. of the cultural relic, and thus it is difficult to appreciate the cultural value of the cultural relic.
[0004] With the development of technology, users can use smartphones to obtain additional introductions of the cultural relic information on the Internet by means of scanning and keyword retrieval. However, in this method, not only is the user operation complex and they need to search for relevant content in a vast amount of resources, but also the display method is often limited to text, images, videos, etc., and is still not vivid enough. Summary of the Invention
[0005] To solve the above problems, this application proposes a method for displaying cultural relics based on XR, including: Scanning and identifying the exhibited cultural relics through an XR device worn by the user to obtain the cultural relic information of the exhibited cultural relics; Determining the social association information of the exhibited cultural relics with the current era according to the cultural relic information; If the degree of association corresponding to the social association information is lower than a preset threshold, then according to the cultural relic information, generating display content for describing the use of the exhibited cultural relics in the historical era; If the degree of association corresponding to the social association information is higher than a preset threshold, then according to the social association information and the cultural relic information, generating display content for describing the use of the exhibited cultural relics in the current era; Displaying the display content to the user through the XR device.
[0006] In one example, determining the social association information of the exhibited cultural relics with the current era according to the cultural relic information specifically includes: Determining the image information and introduction information contained therein according to the cultural relic information; Perform similarity comparison based on the said image information to select the first object that has a similarity in appearance to the exhibition cultural relic higher than a preset threshold and belongs to the current era; Based on the said introduction information, output, through a multimodal large language model, the second object that has the same use and belongs to the current era; Based on the said first object and the said second object, obtain the social association information between the exhibition cultural relic and the current era.
[0007] In one example, based on the said social association information and the said cultural relic information, generate display content for describing the use of the exhibition cultural relic in the current era, specifically including: Generate video content constraint information according to the social association information; the video content constraint information at least includes: constructing a display video according to the usage scenario of the second object in the current era, and respectively replacing the second object in the display video with the first object and the exhibition cultural relic; Generate corresponding prompt words based on the said social association information, the said cultural relic information and the said video content constraint information, and generate, through a multimodal large language model, a display video for describing the use of the exhibition cultural relic in the current era.
[0008] In one example, generate corresponding prompt words based on the said social association information, the said cultural relic information and the said video content constraint information, and generate, through a multimodal large language model, a display video for describing the use of the exhibition cultural relic in the current era, specifically including: Generate a first text prompt word based on the said social association information and the said cultural relic information, and output, through a multimodal large language model, the corresponding first text information; the first text information is used to describe the use of the cultural relic information in the current era; Generate a second text prompt word based on the first text information, and output, through a multimodal large language model, the corresponding second text information; the second text information includes video creation information corresponding to converting the use from the text modality to the video modality; Generate a third text prompt word based on the second text information, the said image information and the said video content constraint information, and output, through a multimodal large language model, the corresponding display video as the display content for describing the use of the exhibition cultural relic in the current era.
[0009] In one example, generate display content for describing the use of the exhibition cultural relic in a historical era, specifically including: Determine the corresponding historical information of the exhibition cultural relic in the historical era according to the said cultural relic information; Generate display content for describing the use of the exhibition cultural relic in the said historical era according to the said historical information; Splice the display content corresponding to the historical eras in chronological order of the historical eras.
[0010] In one example, the method further includes: Determine the image information corresponding to the exhibition cultural relics and the observation information of the XR device according to the cultural relic information; Determine the damaged area of the exhibition cultural relics according to the image information; Virtually repair the exhibition cultural relics according to the damaged area and generate a virtual repair process corresponding to the observation information; Continuously display the exhibition cultural relics through the first area of the XR device, and display the virtual repair process and the exhibition cultural relics obtained after the virtual repair process on the basis of the exhibition cultural relics through the second area of the XR device.
[0011] In one example, the method further includes: Determine multiple XR devices that are currently observing the same exhibition cultural relic according to the positioning information and observation information of the XR device, and form a dynamic device set with the multiple XR devices; If it is determined that there is a specified situation according to the XR device of the current user, select the XR devices of other users that do not have the specified situation in the dynamic device set as shared devices; Share the display content in the shared devices with the XR device of the current user until the specified situation is eliminated.
[0012] In one example, if it is determined that there is a specified situation according to the XR device of the current user, select the XR devices of other users that do not have the specified situation in the dynamic device set as shared devices, which specifically includes: If it is determined that there is an occlusion area on the observation path according to the observation information of the XR device of the current user, and there are occlusion areas on the observation paths of more than a preset proportion of the XR devices in the dynamic device set, it is considered that there is a first specified situation; If the transmission speed between the XR device of the current user and the cloud server is lower than the preset speed, it is considered that there is a second specified situation; For the first specified situation, select the XR device of other users that is closest to the exhibition cultural relic and does not have the first specified situation as the shared device; For the second specified situation, select the XR device of other users that is closest to the current user and does not have the second specified situation as the shared device.
[0013] On the other hand, the present application also proposes an XR-based cultural relic display device, including: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the XR-based cultural relic display method as described in any of the above examples.
[0014] On the other hand, the present application also proposes an XR-based cultural relic display device, including: A scanning module that scans and identifies the exhibited cultural relics through an XR device worn by the user to obtain the cultural relic information of the exhibited cultural relics; An association module that determines the social association information between the exhibited cultural relics and the current era according to the cultural relic information; A first content generation module that, if the degree of association corresponding to the social association information is lower than a preset threshold, generates display content for describing the use of the exhibited cultural relics in the historical era according to the cultural relic information; A display module that displays the display content to the user through the XR device.
[0015] On the other hand, the present application also proposes a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are configured to implement the XR-based cultural relic display method as described in any of the above examples.
[0016] The XR-based cultural relic display method proposed by the present application can bring the following beneficial effects: Compared with the traditional method in which users query by themselves through a smart phone, through the XR device for intelligent scenario reconstruction, there is no need for users to perform cumbersome steps such as scanning and searching with a smart phone, and the display method is no longer limited to forms such as text, images, and videos. While reducing the threshold for understanding cultural relics, it can also improve the interest during the cultural relic display process. When the audience wears an XR device to scan cultural relics, by analyzing the degree of association between the cultural relics and modern life, for low-association cultural relics, a dynamic scene is generated to restore the ancient usage scenario, so as to realize the immersive experience of the user for the cultural relic explanation when browsing in the museum, making the cultural relic display process more vivid and vivid. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings: Figure 1 It is a schematic structural diagram of a system on which the embodiments of the present application can run; Figure 2Schematic flowchart of the XR-based cultural relic display method in the embodiments of the present application; Figure 3 Schematic flowchart of determining social association information in one case in the embodiments of the present application; Figure 4 Schematic flowchart of generating display content when the association degree is relatively low in one case in the embodiments of the present application; Figure 5 Schematic flowchart of virtual restoration of cultural relics in one case in the embodiments of the present application; Figure 6 Schematic flowchart of the display encountering special situations in one case in the embodiments of the present application; Figure 7 Schematic diagram of the module structure of the XR device in one case in the embodiments of the present application; Figure 8 Schematic diagram of the XR-based cultural relic display device in one case in the embodiments of the present application; Figure 9 Schematic diagram of the XR-based cultural relic display device in one case in the embodiments of the present application. Detailed implementation manners
[0018] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0019] The following details the technical solutions provided by each embodiment of the present application in conjunction with the drawings.
[0020] The XR-based cultural relic display method provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. As Figure 1 shown, the application environment may include: the XR device 101 worn by the user 107, the communication network 102, the cloud server 103, the text resource server 104, the image resource server 105, and the database 106.
[0021] Among them, XR stands for Extended Reality, which is a collective term for Virtual Reality (VR), Augmented Reality (AR), and Mixed Reality (MR). XR devices can be set as VR glasses, AR glasses, MR glasses, etc. For the convenience of description, in the embodiments of this article, the XR device is taken as an XR glasses for explanation. The communication network 102 can serve as a data transmission channel to provide a communication link for the communication between the XR device 101 and the cloud server 103. So that after the XR device 101 collects various information such as sound signals, perspective information, and image information, it can be transmitted to the cloud server 103 based on the communication network 102. At the same time, the information processed by the cloud server 103 can also be fed back to the XR device 101 for rendering and display. The cloud server 103 is a server that can provide various services. In order to improve the processing efficiency, the cloud server 103 is connected to the text resource server 104 and the image resource server 105, so that after receiving the information collected by the XR device 101 based on the communication network 102, it can be used to obtain and process the text information and image information related to cultural relics respectively. After the relevant information is processed, the cloud server 103 combines the processing results of the cultural relics-related information of the XR device 101 and feeds it back to the XR device 101 based on the communication network 102 and presents it to the user. Among them, the database 106 is respectively connected to the text resource server 104 and the image resource server 105, and can be used to store and manage the text information of the text resource server 104 and the image information of the image resource server 105. The database 106 can be integrated on the cloud server 103 or placed on other network servers.
[0022] Such as Figure 2 As shown, the embodiments of this application provide a method for displaying cultural relics based on XR, including: S201: Scan and identify the exhibited cultural relics through the XR device worn by the user to obtain the cultural relics information of the exhibited cultural relics.
[0023] Set up a cloud server, which can process the corresponding text information and image information, store data such as 2D images, 3D models, historical data, and material spectrum data of cultural relics through a database, and can interact with XR glasses to achieve corresponding data transmission and processing functions.
[0024] Cultural relic information can include image information of exhibited cultural relics (hereinafter referred to as cultural relics for convenience of description, including 2D images, 3D images, etc.), nameplate information, introduction information, etc. Among them, the introduction information can be obtained through the text introduction beside the cultural relic. At the same time, the introduction information can also be complemented by records in the connected or set cloud database. For example, information such as the age, name, and usage method of the cultural relic can be added.
[0025] At the same time, a knowledge parsing engine can be set up, which can combine a knowledge graph to annotate the names of cultural relic components, craftsmanship techniques, and historical events in real time as cultural relic information.
[0026] Generally speaking, the moving speed of the user can be judged according to the displacement speed of the XR device. When the user is basically stationary, it is considered that the user is interested in the cultural relic in front of him / her. By asking the user through the XR device, the corresponding scanning and recognition process is triggered.
[0027] By scanning cultural relic information, nameplate information, introduction information, etc. through the XR glasses, during the process of the user wearing the XR glasses to visit the exhibition, the cultural relic scanning function of the XR glasses is triggered to obtain relevant information of the cultural relic.
[0028] The user wears XR glasses (for example, the field of view angle is 53°), the acquisition resolution is 1280*800. Through the multi-angle camera of the XR glasses, the user wears the XR glasses, turns on the scanning function of the XR glasses, gazes at the cultural relic and walks around the cultural relic for a circle to complete a 360-degree scan, and collects data such as the surface texture, pattern, and design of the cultural relic; aims at the nameplate of the cultural relic and directly scans the nameplate introduction information through the camera. Of course, if the current number of people is too large or restricted by the venue, only partial degrees can be scanned or only scanned in place to obtain the corresponding two-dimensional image.
[0029] Through sourcing in the full-scale cultural relic information retrieval library, the preliminary retrieval and matching of cultural relic information are completed, and at the same time, digital cultural relic objects of all matching results are generated in the digital space for 360-degree surround display of the search results.
[0030] S202: Determine the social association information between the exhibited cultural relic and the current era according to the cultural relic information.
[0031] The current era can be defined based on human needs. For example, the time after a certain time point (such as a specific year) is defined as the current era, or the time within a certain time period (such as set to 10 years) from the current time point is defined as the current era.
[0032] The social association information refers to the association information between the exhibited cultural relic and the objects still in daily use in the current era. This association information can be in terms of appearance, usage, etc.
[0033] Specifically, as Figure 3 shown, according to the cultural relic information, the included image information and introduction information are determined. Among them, the image information is in the image modality, while the introduction information is in the text modality.
[0034] Based on the image information, similarity comparison is performed to select the first object that has a similarity in appearance to the exhibition cultural relic higher than a preset threshold and belongs to the current era. When performing the comparison, the scope of the comparison library can be the images in a preset daily necessities library or the comparison can be made through open-source image libraries on the Internet. Since large-scale comparisons are usually made during the comparison, a comparison method based on perceptual hashing can be used to compress the image into a hash value of a fixed length and quickly compare the similarity through the Hamming distance.
[0035] Based on the introduction information, a second object that has the same use and belongs to the current era is output through a multimodal large language model. For example, the prompt word at this time can be "You are a senior historical cultural relic expert and are proficient in current life knowledge. Now, here is a cultural relic for you, and its introduction information is {information}. Please find an object with the same use among the objects still used by people in the current era and quantify and score the similarity degree of the uses between the two, with a full score of 100 points." At this time, the object with a score higher than a certain score can be used as the second object. Among them, the content in the {} symbol is a variable and is added based on the actual situation.
[0036] At this time, based on the first object and the second object, the social association information between the exhibition cultural relic and the current era is obtained. For the first object, its appearance is similar to the cultural relic, and for the second object, its use is similar to the cultural relic. Therefore, the cultural relic can be directly described in terms of appearance and use through these two (the first object and the second object can each have multiple, but for the convenience of calculation and description to the user, only one object can be selected from the first object and the second object respectively), and the first object and the second object are displayed in the XR glasses to assist the user in quickly understanding the cultural relic.
[0037] S203: If the association degree corresponding to the social association information is lower than the preset threshold, then according to the cultural relic information, display content for describing the use of the exhibition cultural relic in the historical era is generated.
[0038] Based on the social association information corresponding to the cultural relic information, the association degree between the cultural relic and modern society is judged. The association degree can be obtained by weighted summation of the similarity scores and quantified scores corresponding to the first object and the second object.
[0039] The display content can include forms such as text, images, videos, etc., and its main purpose is to explain the relevant information of the cultural relic to the user.
[0040] If the degree of association is lower than the preset threshold, it indicates that it is difficult for users to quickly understand the cultural relic. At this time, according to the cultural relic information, display content about the usage of the cultural relic in the historical era is generated. For example, corresponding prompting words (including key words, scene descriptions, character settings, action requirements, etc. corresponding to the cultural relic) are generated according to the cultural relic information, and the prompting words are input into the multi-modal large language model to output pictures or videos of the usage scene of the cultural relic in the historical era and display them to help users understand the usage of the cultural relic in the society at that time.
[0041] For example, when a user wears XR glasses and goes to an art museum, the exhibition cultural relic in front of them is a landscape painting. The XR glasses camera scans the landscape painting and then scans the nameplate next to the painting, obtaining that the brushstroke style in the painting is exactly the relevant style advocated by the author. Then, according to the name of the painting in the nameplate, the inspiration source when the author created this painting, the creation process, and the video of the painting's transfer to the museum after creation are generated, telling the story of this painting from its birth to its spread to the present. At this time, the usage of the exhibition cultural relic is not only limited to its use as a painting itself for appreciation, but can also include the story behind the creation.
[0042] Also for example, for the historical introduction of utensils, after identifying the shape, texture, and color of the utensils and combining the museum nameplate information, information such as the existence era, name, and usage method of the cultural relic is obtained, and a video is generated through the large model's ability of text-to-video to explain the corresponding cultural relic.
[0043] Furthermore, as Figure 4 shown, according to the cultural relic information, the corresponding historical information in the historical era is determined. For example, through multi-modal large language models, the Internet, introduction information, etc., the cultural relic information of the cultural relic at each historical time node (including its usage, experienced events, etc. at each historical time node) is obtained.
[0044] According to the historical information, display content for describing the usage of the exhibition cultural relic in the historical era is generated. For each historical era, according to the cultural relic information corresponding to that historical era, prompting words are generated and input into the multi-modal large language model to obtain the node display video of the cultural relic in that historical era.
[0045] In the chronological order of historical eras, the display content corresponding to the historical eras is spliced. After obtaining the display videos of each historical era, the node display videos are spliced in chronological order to obtain a complete display video and play it.
[0046] In another case, if the degree of association corresponding to the social association information is higher than the preset threshold, it is considered that with the cooperation of appropriate display content, the user's understanding of the cultural relic can be accelerated. Therefore, at this time, according to the social association information and the cultural relic information, display content for describing the use of the exhibition cultural relic in the current era is generated.
[0047] Specifically, for cultural relics with a relatively high degree of association corresponding to the social association information, if they are still explained in the above-mentioned way, it may not be vivid enough, and it is still difficult for users to associate the cultural relic with the society in the current era. Therefore, the use of the cultural relic in the current era can be used to generate display content for display, making the entire demonstration process more vivid.
[0048] According to the social association information, video content constraint information is generated. Video content constraint information refers to what constraints need to be imposed when generating the corresponding display video, and this constraint can be added to the corresponding prompt words for generating the display video.
[0049] The video content constraint information at least includes: constructing a display video according to the usage scenario of the second object in the current era, and respectively replacing the second object in the display video with the first object and the exhibition cultural relic.
[0050] Since the use of the second object is basically similar to that of the cultural relic, a display video is first constructed according to it. For example, if the cultural relic is a three-legged jue cup, its main use is for drinking alcohol. At this time, the second object can be a wine glass in the current era, and a drinking scenario is constructed based on the wine glass. And the first object can be a three-legged mug in modern handicrafts.
[0051] When replacing with the first object and the exhibition cultural relic respectively, two sub-videos can be generated. For example, in the first half of the already constructed display video, the user uses the first object to replace the second object, and in the second half of the display video, the user uses the exhibition cultural relic to replace the second object, and then the first half and the second half are spliced to obtain the display video. In this way, in the display video, not only can the use of the cultural relic be vividly displayed to the user through a modern scenario, but also the appearance-similar ancient and modern objects can appear in the same frame, increasing the user's sense of immersion in the scenario.
[0052] At this time, according to the social association information, the cultural relic information, and the video content constraint information, corresponding prompt words are generated, and a display video for describing the use of the exhibition cultural relic in the current era is generated through a multi-modal large language model.
[0053] Furthermore, when generating the prompt, the first text prompt can be generated based on the social association information and the cultural relic information, and the corresponding first text information can be output through the multi-modal large language model; the first text information is used to describe the uses of the cultural relic information in the current era. For example, the first text prompt can be "You are a senior historical cultural relic expert who is also proficient in current life knowledge. Now, here is a cultural relic with its introduction information {information}, and its social associations with the current society include {relevance}. You are responsible for associating based on the social associations and imagining what similar uses this cultural relic can have in which scenarios if it is placed in the current society."
[0054] Generate the second text prompt based on the first text information, and output the corresponding second text information through the multi-modal large language model; the second text information includes the video creation information corresponding to the conversion of the use from the text modality to the video modality. For example, the second text prompt can be "You are a senior historical cultural relic expert who is also proficient in current life knowledge and is an expert proficient in photography and video creation. Now, here is a text information {use} for describing the uses of the cultural relic in the current era. You are required to generate a video creation information for video creation, and the video creation information includes but is not limited to: main line structure, storyboard design, visual style, sound requirements, technical constraints."
[0055] Generate the third text prompt based on the second text information and the image information, and output the corresponding display video through the multi-modal large language model as the display content for describing the uses of the exhibition cultural relic in the current era. For example, the third text prompt can be "You are a senior video creator. Here are the relevant requirements {create} corresponding to the video creation information, here is the image information {image} of the cultural relic in the video creation, and here is the video content constraint information {constraint}. You are required to generate a video based on the video creation information and the image information to display the uses of this cultural relic in the current society, and at the same time, it is necessary to meet the requirements in the video content constraint information."
[0056] S204: Display the said display content to the said user through the said XR device.
[0057] If the user has questions during the display process and can ask questions related to the cultural relic through voice, then input the asked question and the cultural relic information into the multi-modal large language model together. For example, for the landscape painting mentioned above, the user asks "What creative techniques are used in this work?" After receiving the question, the large model combines the relevant introduction information recognized by XR scanning and generates the corresponding display video to explain to the user. At this time, "Video is loading..." is displayed.
[0058] During the loading process of the display video, the display video is generated by inputting the large model through the split-shot content design. After the generation is completed, the loading state is converted to the video screen image, and after displaying "Start automatic play", the video display begins.
[0059] Through intelligent context reconstruction, the threshold for understanding cultural relics is lowered. When visitors wear XR devices to scan cultural relics, the relevance of the cultural relics to modern life is analyzed. For highly relevant cultural relics, the comparison data with modern products is superimposed in real time; for low-relevance cultural relics, dynamic scenes are generated to restore the ancient usage context, enabling users to immerse themselves in the explanation of cultural relics when browsing in the museum and increase the fun.
[0060] In one embodiment, virtual restoration can also be performed through XR to show users the complete appearance of the cultural relics.
[0061] Specifically, Figure 5 As shown, the image information corresponding to the exhibited artifacts and the observation information of the XR device are determined based on the artifact information. A high-precision 3D model of the artifact can be constructed through laser scanning and multispectral imaging using XR devices or laser equipment in advance, and the material, texture, and historical context information can be stored in the cloud. The observation information can include the observation angle and observation distance of the exhibited artifacts.
[0062] Determine the defective area of the exhibited cultural relics based on the image information. For example, a lightweight model is set up on the XR glasses to detect the defective area in real time. The lightweight model backbone network can select MobileNetV3-Small, and set the corresponding detection head based on YOLOv5s (including a three-dimensional convolution layer, a depth-separable convolution layer, and a lightweight attention mechanism) to detect the defective area and output the coordinates and type of the defective area.
[0063] Virtually repair the exhibited cultural relics according to the damaged areas, and generate a virtual repair process corresponding to the observed information. For example, based on the Generative Adversarial Networks (GAN) and matching with historical data, a repaired virtual 3D model is generated. Alternatively, through a multimodal large language model, the exhibited cultural relics marked with damaged areas are input, and instructions are issued to generate a virtual repair process under the current observation information (including observation angle and observation distance). At this time, the prompt word can be "You are a senior cultural relic restoration expert. The image information of the current cultural relic is {image}, in which the defective area has been marked. Use your professional knowledge to generate a video corresponding to the repair process of the cultural relic under the current observation information {observe}."
[0064] The content of virtual restoration can be targeted at damaged artifacts or paintings and calligraphy. After identifying the damaged area, a virtual restoration process is generated through completion. It can also be targeted at faded murals. When the faded area is identified, the original color simulation effect is superimposed as the virtual restoration process.
[0065] At this time, through the first area of the XR device, the exhibition artifacts are continuously displayed, and through the second area of the XR device, based on the exhibition artifacts, the virtual restoration process and the exhibition artifacts obtained after the virtual restoration process are displayed.
[0066] Among them, after the virtual restoration process is generated, the virtual restoration model can be spatially aligned with the real artifact through SLAM technology, which can support the seamless visual fusion of the virtual artifact and the real artifact, and at the same time can also support the virtual artifact to be displayed in space against the real artifact.
[0067] When the XR device is an XR glasses, one lens can be used as the first area and the other lens can be used as the second area. At this time, the current appearance of the artifact and the restored appearance are displayed simultaneously, improving the user's interest in artifact restoration. And since the restoration process is generated based on the current observation information, the content seen in the two lenses is basically the same and will not affect the user's viewing.
[0068] In one embodiment, when visiting places such as museums, there may be some special situations that make it difficult for the XR device to display the content normally.
[0069] Based on this, as Figure 6 shown, according to the positioning information and observation information of the XR device, multiple XR devices that are currently observing the same exhibition artifact are determined, and these multiple XR devices are formed into a dynamic device set. According to the positioning information, the XR devices in the same area are determined, and then according to the observation information, the XR devices that observe the same artifact among them are determined.
[0070] The XR devices in the dynamic device set are dynamic. Each exhibition artifact corresponds to a dynamic device set. When the user goes to see other artifacts, they automatically leave the dynamic device set corresponding to the current exhibition artifact.
[0071] At this time, if it is determined according to the current user's XR device that there is a specified situation, then in the dynamic device set, other users' XR devices that do not have the specified situation are selected as shared devices. The display content in the shared devices is shared with the current user's XR device until the specified situation is eliminated.
[0072] Sharing the display content for the current user through the shared device enables the current user's XR device to quickly avoid when encountering the specified situation and will not affect the current user's observation experience.
[0073] Furthermore, for the observation situation, if it is determined according to the observation information of the XR device of the current user that there is an occlusion area on the observation path, and there is an occlusion area on the observation paths of more than a preset proportion of XR devices in the dynamic device set, it is considered that there is a first specified situation.
[0074] If only the current user has the problem of occlusion area, the user can be reminded to move slightly. However, if multiple users have occlusion problems, it is considered that the proportion range of the occlusion area is too large due to the large number of people visiting currently.
[0075] Meanwhile, if the transmission speed between the XR device of the current user and the cloud server is lower than the preset speed, it is considered that there is a second specified situation. At this time, it may be caused by an abnormality in the current environment or the transmission module of the device.
[0076] At this time, for the first specified situation, select the XR device of another user that is closest to the exhibition cultural relic and does not have the first specified situation as the shared device. This shared device is very close to the exhibition cultural relic and can completely observe the shape of the exhibition cultural relic. Therefore, it is selected as the shared device to share the display content.
[0077] For the second specified situation, select the XR device of another user that is closest to the current user and does not have the second specified situation as the shared device. For the XR device with poor transmission speed, select the closest XR device as the shared device, and then establish D2D communication between this shared device and the current XR device (for example, communicate through Wi-Fi Direct / Ultra Wideband), which can ensure that the current XR device can normally play the display content and improve the user experience.
[0078] The embodiment of the present application also provides an XR device, as Figure 7 shown, the XR device can be an XR glasses, including a processor 701, a memory 702, a shooting module 703, a display module 704, an interaction module 705, and a wireless communication module 706.
[0079] The memory 702 can be used to store computer programs, for example, software programs and modules of application software. The processor 701 executes various functional applications and data processing by running the computer programs stored in the memory 702, such as implementing the method provided by the embodiment of the present application.
[0080] The shooting module 703 is mainly responsible for shooting the introduction information and image information of the cultural relic, and the shooting module can be a camera.
[0081] The display module 704 is used to display the cultural relic restoration result obtained by the method provided by the embodiment of the present application, etc.
[0082] The interaction module 705 includes a microphone array, etc., which is used to implement the interaction between the XR glasses and the user, facilitating the user to input instructions.
[0083] The wireless communication module 706 includes: 4G / 5G full-network communication, WiFi, Bluetooth, etc.
[0084] Those of ordinary skill in the art can understand that Figure 7 the structure shown is only schematic and does not limit the structure of the XR glasses. For example, the XR glasses may further include more or fewer components than those shown Figure 7 in the figure, or have a different configuration from that shown Figure 7 in the figure.
[0085] The following is a brief introduction to the application scenarios of the XR glasses.
[0086] During the use of the XR glasses, when the user wears the XR glasses to visit a museum, based on the user's trigger or pre-setting, the XR glasses scan the cultural relic information of the exhibited cultural relics, including introduction information, image information, etc., and transmit the cultural relic information to the cloud server. The cloud server determines the social association information of the exhibited cultural relics with the current era, and generates different display contents based on the level of the social association information, and feeds them back to the XR glasses. The XR glasses display the display contents to the user, so that the user can understand the exhibited cultural relics more vividly.
[0087] As Figure 8 shown, the embodiment of the present application also provides an XR-based cultural relic display device, including: A scanning module 801 scans and identifies the exhibited cultural relics through the XR device worn by the user to obtain the cultural relic information of the exhibited cultural relics; An association module 802 determines the social association information of the exhibited cultural relics with the current era according to the cultural relic information; A first content generation module 803, if the association degree corresponding to the social association information is lower than a preset threshold, generates a display content for describing the use of the exhibited cultural relics in the historical era according to the cultural relic information; A display module 804 displays the display content to the user through the XR device.
[0088] Optionally, the association module 802 determines the image information and introduction information included therein according to the cultural relic information; Performs a similarity comparison according to the image information to select a first object that has a similarity to the appearance of the exhibited cultural relic higher than a preset threshold and belongs to the current era; According to the introduction information, outputs a second object that has the same use and belongs to the current era through a multimodal large language model; According to the first object and the second object, social correlation information between the exhibited cultural relics and the current era is obtained.
[0089] Optionally, the method further includes: a second content generation module 805, which generates video content constraint information according to the social association information if the association degree corresponding to the social association information is higher than a preset threshold; the video content constraint information at least includes: constructing a display video according to the usage scenario of the second object in the current era, and replacing the second object in the display video with the first object and the exhibition cultural relic respectively; Corresponding prompt words are generated according to the social association information, the cultural relic information and the video content constraint information, and a display video for describing the use of the exhibited cultural relics in the current era is generated through a multimodal large language model.
[0090] Optionally, the second content generation module 805 generates a first text prompt word according to the social association information and the cultural relic information, and outputs corresponding first text information through a multimodal large language model; the first text information is used to describe the purpose of the cultural relic information in the current era; Generate a second text prompt word according to the first text information, and output the corresponding second text information through a multimodal large language model; the second text information includes video creation information corresponding to the conversion of the purpose from text mode to video mode; A third text prompt word is generated according to the second text information, the image information, and the video content constraint information, and a corresponding display video is output through a multimodal large language model as display content describing the use of the exhibited cultural relics in the current era.
[0091] Optionally, the first content generation module 803 determines the historical information corresponding to the cultural relic in the historical era according to the cultural relic information; Generating display content for describing the use of the exhibited cultural relics in the historical era according to the historical information; According to the chronological order of the historical eras, the display contents corresponding to the historical eras are spliced together.
[0092] Optionally, it further includes: a virtual repair module 806, which determines the image information corresponding to the exhibition cultural relic and the observation information of the XR device according to the cultural relic information; Determine the defective area of the exhibited cultural relic according to the image information; Virtually repairing the exhibited cultural relic according to the damaged area, and generating a virtual repair process corresponding to the observation information; Continuously display the exhibition cultural relics through the first area of the XR device, and through the second area of the XR device, on the basis of the exhibition cultural relics, display the virtual restoration process and the exhibition cultural relics obtained after the virtual restoration process.
[0093] Optionally, it further includes: a device sharing module 807, which determines multiple XR devices currently observing the same exhibition cultural relic according to the positioning information and observation information of the XR device, and forms a dynamic device set with these multiple XR devices; If it is determined that there is a specified situation according to the XR device of the current user, then among the dynamic device set, select the XR devices of other users that do not have the specified situation as shared devices; Share the display content in the shared devices with the XR device of the current user until the specified situation is eliminated.
[0094] Optionally, for the device sharing module 807, if it is determined according to the observation information of the XR device of the current user that there is an occlusion area on the observation path, and there is an occlusion area on the observation paths of more than a preset proportion of the XR devices in the dynamic device set, then it is considered that there is a first specified situation; If the transmission speed between the XR device of the current user and the cloud server is lower than the preset speed, then it is considered that there is a second specified situation; For the first specified situation, select the XR device of other users that is closest to the exhibition cultural relic and does not have the first specified situation as the shared device; For the second specified situation, select the XR device of other users that is closest to the current user and does not have the second specified situation as the shared device.
[0095] As Figure 9 shown, an embodiment of the present application further provides an XR-based cultural relic display device, including: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the XR-based cultural relic display method as described in any of the above embodiments.
[0096] An embodiment of the present application further provides a non-volatile computer storage medium, storing computer-executable instructions, and the computer-executable instructions are set to implement the XR-based cultural relic display method as described in any of the above embodiments.
[0097] The embodiments in the present application are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for relevant content.
[0098] The devices and media provided in the embodiments of the present application correspond one-to-one with the methods. Therefore, the devices and media also have beneficial technical effects similar to those of their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.
[0099] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. An XR-based cultural relic display method, characterized in that, Including: Scanning and identifying exhibition cultural relics through an XR device worn by a user to obtain the cultural relic information of the exhibition cultural relics; Determining the social association information between the exhibition cultural relics and the current era according to the cultural relic information; If the association degree corresponding to the social association information is lower than a preset threshold, generating display content for describing the use of the exhibition cultural relics in the historical era according to the cultural relic information; Displaying the display content to the user through the XR device.
2. The XR-based cultural relic display method according to claim 1, wherein Determining the social association information between the exhibition cultural relics and the current era according to the cultural relic information specifically includes: Determining the image information and introduction information contained therein according to the cultural relic information; Performing similarity comparison according to the image information to select a first object that is similar in appearance to the exhibition cultural relics and belongs to the current era and has a similarity higher than the preset threshold; Outputting a second object that has the same use and belongs to the current era through a multimodal large language model according to the introduction information; Obtaining the social association information between the exhibition cultural relics and the current era according to the first object and the second object.
3. The XR-based cultural relic display method according to claim 2, wherein, After determining the social association information between the exhibition cultural relics and the current era according to the cultural relic information, the method further includes: If the association degree corresponding to the social association information is higher than the preset threshold, generating video content constraint information according to the social association information; the video content constraint information at least includes: constructing a display video according to the use scenario of the second object in the current era, and respectively replacing the second object in the display video with the first object and the exhibition cultural relics; Generating corresponding prompt words according to the social association information, the cultural relic information and the video content constraint information, and generating a display video for describing the use of the exhibition cultural relics in the current era through a multimodal large language model.
4. The XR-based cultural relic display method according to claim 3, wherein, Generating corresponding prompt words according to the social association information, the cultural relic information and the video content constraint information, and generating a display video for describing the use of the exhibition cultural relics in the current era through a multimodal large language model specifically includes: Generating a first text prompt word according to the social association information and the cultural relic information, and outputting corresponding first text information through a multimodal large language model; the first text information is used to describe the use of the exhibition cultural relics in the current era; Generating a second text prompt word according to the first text information, and outputting corresponding second text information through a multimodal large language model; the second text information includes video creation information corresponding to converting the use of the exhibition cultural relics in the current era from text modality to video modality; Generating a third text prompt word according to the second text information, the image information and the video content constraint information, and outputting a corresponding display video through a multimodal large language model as display content for describing the use of the exhibition cultural relics in the current era.
5. The XR-based cultural relic display method according to claim 1, wherein Generating display content for describing the use of the exhibition cultural relics in the historical era according to the cultural relic information specifically includes: Determining the historical information corresponding thereto in the historical era according to the cultural relic information; Generate display content for describing the use of the exhibition cultural relics in the historical era according to the historical information; Splice the display content corresponding to the historical era in chronological order of the historical era.
6. The XR-based cultural relic display method according to claim 1, wherein The method further includes: Determine the image information corresponding to the exhibition cultural relics and the observation information of the XR device according to the cultural relics information; Determine the damaged area of the exhibition cultural relics according to the image information; Virtually repair the exhibition cultural relics according to the damaged area, and generate a virtual repair process corresponding to the observation information; Continuously display the exhibition cultural relics through the first area of the XR device, and display the virtual repair process and the exhibition cultural relics obtained after the virtual repair process on the basis of the exhibition cultural relics through the second area of the XR device.
7. The XR-based cultural relic display method according to claim 1, wherein The method further includes: Determine multiple XR devices that are currently observing the same exhibition cultural relics according to the positioning information and observation information of the XR device, and form a dynamic device set with the multiple XR devices; If it is determined according to the XR device of the current user that there is a specified situation, select the XR devices of other users that do not have the specified situation in the dynamic device set as shared devices; Share the display content in the shared devices with the XR device of the current user until the specified situation is eliminated.
8. The XR-based cultural relic display method according to claim 7, wherein If it is determined according to the XR device of the current user that there is a specified situation, select the XR devices of other users that do not have the specified situation in the dynamic device set as shared devices, specifically including: If it is determined according to the observation information of the XR device of the current user that there is an occlusion area on the observation path, and there is an occlusion area on the observation paths of more than a preset proportion of the XR devices in the dynamic device set, it is considered that there is a first specified situation; If the transmission speed between the XR device of the current user and the cloud server is lower than the preset speed, it is considered that there is a second specified situation; For the first specified situation, select the XR device of other users that is closest to the exhibition cultural relics and does not have the first specified situation as the shared device; For the second specified situation, select the XR device of other users that is closest to the current user and does not have the second specified situation as the shared device.
9. An XR-based cultural relic display device, characterized in that Includes: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the XR-based cultural relic display method according to any one of claims 1 to 8.
10. An XR-based cultural relic display device, characterized in that, Includes: A scanning module that scans and identifies exhibition cultural relics through an XR device worn by a user to obtain the cultural relics information of the exhibition cultural relics; An association module that determines the social association information between the exhibition cultural relics and the current era according to the cultural relics information; A first content generation module that, if the association degree corresponding to the social association information is lower than a preset threshold, generates display content for describing the use of the exhibition cultural relics in the historical era according to the cultural relics information; A display module that displays the said display content to the said user through the said XR device.
Citation Information
Patent Citations
Multi-modal recognition algorithm based on images and characters
CN119380346A
Method for constructing cultural relic knowledge organization, expression and characterization model for Song charm themes
CN120146164A
Educational culture multi-mode interactive learning system and method based on XR augmented reality
CN120161945A
Information presentation system, device, method, and program
WO2023067715A1
Method for generating information, method for displaying information, device, and storage medium
WO2025113666A1