Binocular vision-based AI intelligent shooting system for cultural relic exhibition hall
Through the AI smart shooting system of cultural and museum exhibition halls based on binocular vision, multimodal technology is integrated to generate high-precision three-dimensional models of cultural relics, which solves the problems of single perspective and insufficient interaction in traditional cultural and museum exhibition halls, and realizes the intelligent upgrade of digital protection of cultural relics and cultural communication.
Patent Information
- Application Number
- CN202511140820.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-15
AI Technical Summary
The shooting and display technology of traditional cultural and museum exhibition halls has problems such as single perspective, incomplete information coverage, high cost, low efficiency, difficulty in supporting real-time interaction and difficulty in data retrieval, resulting in low audience participation and inconvenient resource management.
The AI smart shooting system for cultural and museum exhibition halls based on binocular vision is adopted, including a binocular vision image acquisition module, an AI image stitching and fusion module, a cultural relic 3D reconstruction module and a real-time fusion smart shooting module. Combined with a multi-strategy trigger mechanism, it generates high-precision 3D models of cultural relics and supports real-time interaction.
It has achieved an intelligent upgrade of cultural relics display and visitor interaction, generated high-precision three-dimensional models of cultural relics, enhanced visitors' immersion and participation, and provided convenient data management and sharing functions.
Smart Images

Figure CN120726575A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cultural relics digitization and interactive technology, and specifically to an AI smart photography system for cultural and museum exhibition halls based on binocular vision. Background Art
[0002] With the vigorous development of the cultural tourism industry, cultural and museum exhibition halls, as important carriers of cultural heritage protection and dissemination, have an increasing demand for display methods and audience interactive experience. Traditional exhibition methods mostly rely on static objects and graphic descriptions, which makes it difficult to fully present the details of cultural relics and historical background, especially in the three-dimensional morphology and texture feature dimensions. There are display limitations. At the same time, breakthroughs in computer vision, artificial intelligence and three-dimensional reconstruction technology have provided new solutions for the digitization of cultural relics. Stereoscopic imaging technology based on binocular vision can achieve multi-perspective high-precision image acquisition by simulating the principle of human eye parallax; image fusion and three-dimensional modeling technology combined with deep learning algorithms can further generate realistic three-dimensional models. Therefore, how to integrate multimodal technology into cultural and museum scenes has become a key issue to improve exhibition effects and audience participation.
[0003] The shooting and display technology of traditional cultural and museum exhibition halls mainly relies on monocular cameras or fixed-angle photography equipment, which has the problems of single perspective and incomplete information coverage, making it difficult to fully present the complex structure and detailed texture of cultural relics. For example, the flat image taken by a monocular camera cannot reflect the three-dimensional appearance of the cultural relics, and manual three-dimensional reconstruction requires professional equipment and complex operating procedures, which is costly and inefficient. In addition, traditional technologies are usually difficult to support real-time interactive functions, resulting in limited audience participation and the inability to generate personalized content through action interaction, resulting in serious homogeneity of the visiting experience. In terms of resource management, traditional systems mostly use distributed storage, and the cultural relic data has a single dimension and is difficult to call, which makes it difficult to meet the needs of cross-platform sharing and secondary development. These limitations restrict the digital transformation and innovative development of cultural and museum exhibition halls. Summary of the Invention
[0004] The purpose of this invention is to make up for the shortcomings of the existing technology and provide an AI smart shooting system for cultural relics exhibition halls based on binocular vision, which includes a binocular vision image acquisition module that can shoot cultural relics in multiple directions; an AI image stitching and fusion module that can denoise, correct and align fused images to generate panoramic images; a cultural relic three-dimensional reconstruction module that constructs a three-dimensional grid model and generates a three-dimensional model with fine texture after optimization; a real-time fusion smart shooting module that fuses the three-dimensional model with the visitor image to generate an interactive picture; an interactive output and management module that provides a preview function and classifies and stores data.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: an AI smart shooting system for cultural and museum exhibition halls based on binocular vision, which includes: a binocular vision image acquisition module, an AI image splicing and fusion module, a cultural relic 3D reconstruction module, a real-time fusion smart shooting module, and an interactive output and management module; The binocular vision image acquisition module uses a binocular camera, which is installed and debugged according to the layout of the cultural relics display cabinet and the characteristics of the cultural relics. It shoots cultural relics from multiple directions and angles according to preset strategies, and obtains and stores high-definition original images of cultural relics from different perspectives; The AI image stitching and fusion module receives binocularly captured multi-view high-definition images and performs denoising and color correction preprocessing. It registers the images using a multi-scale feature flow matching algorithm and fuses the registered images with the help of a depth perception fusion function to generate a panoramic image of the cultural relics. The cultural relic 3D reconstruction module: builds a 3D grid model based on the image data of the binocular vision image acquisition module, introduces the panoramic image of the AI image stitching and fusion module, and generates a 3D model of the cultural relic with fine texture through texture and structure symbiotic optimization; The real-time fusion intelligent shooting module retrieves the 3D model of the cultural relics 3D reconstruction module, collects and pre-processes the real-time image of the visitor, fuses the two through a dynamic interactive anchor point fusion algorithm, and simultaneously recognizes the visitor's movements to generate an interactive photo screen; The interactive output and management module receives the photo results of the real-time fusion smart shooting module, provides preview, save and sharing functions, stores cultural relics data by dimension classification and opens interfaces for calling.
[0006] Furthermore, in the binocular vision image acquisition module, preset strategies include a timing trigger strategy and an induction trigger strategy; The timing trigger strategy is to start the binocular camera shooting at a fixed time interval. The time interval can be dynamically adjusted within the range of 1 minute to 10 minutes through the exhibition hall management terminal. After adjustment, the system completes the strategy update within 30 seconds; The induction triggering strategy is: using an array of infrared sensors to detect the state of a human body within a range of 1 to 3 meters in front of the cultural relics display cabinet, and triggering the binocular camera to shoot when it is recognized that the human body stays for ≥5 seconds.
[0007] Furthermore, in the AI image stitching and fusion module, a multi-scale feature flow matching algorithm is used to align images, and the calculation formula is: ,in are the left and right view feature points, is the scale weight, Control the degree of penalty for flow field differences, and Both are left and right perspectives The feature vector of the scale, is the L2 norm.
[0008] Furthermore, in the AI image stitching and fusion module, the registered images are fused with the help of the depth perception fusion function, and the calculation formula is: ,in The fused image is The pixel value at and Both are images after left and right view registration. is the fusion weight, and the calculation formula is: , and Both left and right perspective images are in The depth value at is the depth difference threshold, is the hyperbolic tangent function.
[0009] Furthermore, in the cultural relic three-dimensional reconstruction module, the specific steps of constructing a three-dimensional grid model based on the image data of the binocular vision image acquisition module are as follows: based on the image data obtained by the binocular vision image acquisition module, the binocular stereo vision principle is used to calculate the parallax between images of different perspectives, and the parallax generated by multi-angle shooting is dynamically adjusted through dynamic parallax compensation related calculations. Then, based on the triangulation principle, the depth information of the surface points of the cultural relic is obtained in combination with the adjusted parallax, and the point cloud data of the cultural relic is constructed. The curvature-guided triangulated model is used to generate triangular facets, and the grid accuracy is adjusted to form a three-dimensional grid model that adapts to the multi-scale characteristics of the cultural relic.
[0010] Furthermore, in the cultural relic 3D reconstruction module, the calculation formula for the parallax between images from different perspectives after compensation is: ,in, It is Collection time After compensation, the parallax at is the original disparity, For the The viewing angle offset of the acquisition, and is the compensation coefficient, which controls the effect of the offset on the compensation result. is the gradient of the original disparity.
[0011] Furthermore, in the cultural relic 3D reconstruction module, the curvature-guided triangulated model is combined to generate triangular facets, and the calculation formula is: ,in It is a three-point group The triangulated probability of is the curvature value corresponding to the three points, and the probability of triangulation of the three-point group is calculated. is the curvature penalty coefficient, is the Euclidean distance, the screening probability The vertex group generates triangles.
[0012] Furthermore, in the cultural relic 3D reconstruction module, a 3D model of a cultural relic with fine texture is generated through texture structure symbiosis optimization, and the specific steps are as follows: (1) Obtain the panoramic image of cultural relics generated by the AI image stitching and fusion module as the original material for texture mapping; (2) Extract the structural features corresponding to each vertex on the surface of the 3D mesh model and determine the curvature properties of different areas on the model surface; (3) Perform texture segmentation on the panoramic image to divide the texture segments corresponding to the different curvature areas of the 3D mesh model; (4) Map the segmented texture segments to the corresponding areas of the 3D mesh model so that the texture fits the surface structure of the model; (5) Adjust the texture fit on the model surface, reduce texture stretching in areas with curvature values K>0.8, and ensure texture integrity in areas with curvature values K≤0.8. The curvature value calculation is as follows: for the triangulated mesh vertices of the three-dimensional model of the cultural relics, select the 1-ring neighborhood vertices to form a local area; by analyzing the surface morphology of the local area, obtain a quantitative index reflecting the surface curvature characteristics as the curvature value of the vertex, which is used to measure the curvature degree of the model surface; the curvature threshold is determined: select cultural relics exhibits to construct a test model, and after systematic three-dimensional reconstruction, analyze the texture stretching deformation under different curvature values; when the curvature value exceeds a certain critical value, the texture stretching rate increases significantly, and the critical value is defined as the curvature threshold; It has been verified that when the curvature value is greater than 0.8, the texture stretching risk meets the need to reduce stretching judgment, and when the curvature value is less than or equal to 0.8, the texture fit meets the requirement of ensuring integrity, so 0.8 is used to distinguish between areas with larger curvature and areas with smaller curvature; (6) After completing the texture mapping of all areas, integrate them to form a three-dimensional model of the cultural relic with fine textures.
[0013] Furthermore, in the real-time fusion smart shooting module, the calculation formula of the dynamic interaction anchor point fusion algorithm is: ,in, is the real-time coordinate of the visitor's hand key point in the camera coordinate system, The coordinates of the interaction anchor point of the cultural relic 3D model are, It is the coordinates of the real-time coordinates of the visitor's hand key points in the camera coordinate system and the coordinates of the interactive anchor points of the cultural relic 3D model. is the Euclidean distance between the hand and the initial position of the anchor point, The default position of the anchor point in a static state is the fusion sharpness parameter, is the dynamic scaling factor, is the maximum interaction distance threshold.
[0014] Compared with existing technologies, this binocular vision-based AI smart photography system for cultural and museum exhibition halls has the following beneficial effects: 1. The present invention realizes the intelligent upgrade of cultural relic display and visitor interaction by integrating binocular vision image acquisition, AI image stitching and fusion, cultural relic 3D reconstruction and real-time fusion smart shooting technology modules. The system adopts a binocular camera shooting mechanism with multi-strategy triggering, combined with timing and induction triggering modes, which can flexibly adapt to the needs of different exhibition scenes, ensuring the high efficiency of cultural relic image acquisition and avoiding resource waste. Through multi-scale feature flow matching and depth perception fusion algorithm, the system can generate high-precision, low-distortion panoramic images and 3D models of cultural relics, significantly improving the problems of single perspective and easy loss of details in traditional shooting, and providing richer data support for the digital protection and research of cultural relics.
[0015] 2. Through the real-time fusion of the smart shooting module, the present invention can collect visitor images and dynamically integrate them with the three-dimensional model of cultural relics. Combined with motion recognition technology, personalized interactive photo images are generated, breaking through the limitations of traditional static displays and enhancing the immersion and participation of visitors. At the same time, the interactive output and management module provides a one-stop service of preview, saving, and sharing, and supports the classification storage and interface call of cultural relics data, providing a convenient solution for the digital operation and cultural communication of the exhibition hall.
[0016] Other advantages, objects and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art based on an examination of the following or may be learned from the practice of the invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.
[0018] Figure 1 This is a flowchart of the AI smart photography system for cultural and museum exhibition halls based on binocular vision; Figure 2 This is a data flow framework diagram of the AI smart shooting system for cultural and museum exhibition halls based on binocular vision. DETAILED DESCRIPTION
[0019] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.
[0020] Example 1: Interactive experience scene in the calligraphy and painting exhibition hall.
[0021] A calligraphy and painting exhibition hall has introduced a binocular vision-based AI smart photography system for cultural and museum exhibition halls. The system aims to enable visitors to interact immersively with ancient paintings through digital technology, while also digitally recording and protecting the paintings. The specific process is as follows: Binocular vision image acquisition module: According to the hanging height of the ancient painting display cabinet, the size of the ancient paintings and the distribution of picture details, the staff installed and debugged binocular cameras 1.5 meters in front of the display cabinet and at a 45° angle on both sides. The system preset a timing trigger strategy and a sensor trigger strategy: the timing trigger strategy automatically shoots at intervals of 3 minutes (the interval can be adjusted within 1-10 minutes through the management terminal and takes effect within 30 seconds after the adjustment), ensuring that the status of the ancient paintings at different time periods can be recorded; an array infrared sensor is deployed within 1-3 meters in front of the display cabinet. When a visitor stays in front of the display cabinet for ≥5 seconds (indicating that he has a strong interest in the ancient paintings), the sensor trigger strategy immediately triggers the camera to shoot. The binocular camera shoots from multiple directions and angles, and finally obtains and stores 20 sets of high-definition original images of the ancient paintings from different perspectives, such as Figure 1 shown.
[0022] AI image stitching and fusion module After receiving multi-view images, we first perform denoising (removing light spots and camera sensor noise caused by light reflection during shooting) and color correction (unifying the brightness and color temperature of images from different viewpoints to restore the original ink shades and color levels of ancient paintings). Then, we use multi-scale feature flow matching to align the images. The calculation formula of the multi-scale feature flow matching algorithm is: ,in are the left and right view feature points, is the scale weight, Control the degree of penalty for flow field differences, and Both are left and right perspectives The feature vector of the scale, It is the L2 norm. By identifying the edge lines of ancient paintings and the outline features of seals, images from different perspectives are accurately aligned in space (for example, ensuring that the edge of the frame from the left perspective is seamlessly connected to the edge of the frame from the front perspective). The depth perception fusion function is then used to fuse the aligned images. The calculation formula of the depth perception fusion function is: ,in The fused image is The pixel value at and Both are images after left and right view registration. is the fusion weight, and the calculation formula is: , and Both left and right perspective images are in The depth value at is the depth difference threshold, It is a hyperbolic tangent function, which ultimately generates a complete panoramic view of the ancient painting, in which the handwriting of the inscriptions and the brushstroke details of the landscape are clearly visible.
[0023] Cultural Relics 3D Reconstruction Module: First, the image data of the binocular vision image acquisition module is used to construct a three-dimensional grid model. The parallax between images of different perspectives is calculated based on the binocular stereo vision principle. The calculation formula is: ,in, It is Collection time After compensation, the parallax at is the original disparity, For the The viewing angle offset of the acquisition, and is the compensation coefficient, which controls the effect of the offset on the compensation result. It is the gradient of the original parallax. The parallax deviation caused by the difference in shooting angle is eliminated through dynamic parallax compensation related calculations. The depth information of the surface points of the ancient painting (including the frame) is obtained by combining the principle of triangulation to construct point cloud data. The curvature-guided triangulation model is then combined to generate triangular facets (screening the vertex group with the triangulation probability that meets the standard). The calculation formula of the curvature-guided triangulation model is: ,in It is a three-point group The triangulated probability of is the curvature value corresponding to the three points, and the probability of triangulation of the three-point group is calculated. is the curvature penalty coefficient, is the Euclidean distance, the screening probability The vertex group is used to generate triangular faces. After adjusting the mesh accuracy, a three-dimensional mesh model that can reflect the three-dimensional sense of the picture frame and the flatness of the paper is formed. Then, the panoramic image of the AI image stitching and fusion module is introduced to extract the structural features of the surface vertices of the three-dimensional mesh model (such as the edges and corners of the picture frame and the curvature of the paper folds). The panoramic image is texture segmented (the texture of the picture frame and the texture of the main body of the picture are separated), and the texture fragments are mapped to the corresponding areas of the model. The texture stretching is reduced in the areas with larger curvature of the corners of the picture frame (to avoid deformation of the wood grain of the picture frame), and the texture is guaranteed to be complete in the areas with small curvature of the main body of the picture (to ensure the continuity of the landscape pattern). Finally, a three-dimensional model of the ancient painting with fine texture is generated.
[0024] Real-time fusion smart shooting module: When a visitor stands in the designated interactive area of the exhibition hall (with a camera in front), the real-time fusion smart shooting module collects the visitor's real-time image through the camera (and performs preprocessing, such as removing background noise and adjusting the portrait brightness). The system then retrieves the 3D model of the ancient painting generated by the cultural relics 3D reconstruction module and fuses the visitor's image with the 3D model of the ancient painting using the dynamic interactive anchor fusion algorithm. The calculation formula of the dynamic interactive anchor fusion algorithm is: ,in, is the real-time coordinate of the visitor's hand key point in the camera coordinate system, The coordinates of the interaction anchor point of the cultural relic 3D model are, It is the coordinates of the real-time coordinates of the visitor's hand key points in the camera coordinate system and the coordinates of the interactive anchor points of the cultural relic 3D model. is the Euclidean distance between the hand and the initial position of the anchor point, The default position of the anchor point in a static state is the fusion sharpness parameter, is the dynamic scaling factor, The maximum interaction distance threshold is set. For example, when a visitor stretches out his hand, the system recognizes the hand movement and merges the hand image with the interaction anchor point of the "scroll" in the ancient painting to form the visual effect of "holding a scroll"; if the visitor makes a bowing gesture, the system will generate an interactive picture of "worshiping the ancient people in the painting".
[0025] Interactive output and management module: After receiving the photo results from the real-time fusion smart shooting module, a preview is provided on the display screen in the interactive area (allowing visitors to view the photo effects). It also supports scanning codes to save (saving photos to mobile phone albums) and sharing on social platforms (generating sharing links with the exhibition hall logo). The system classifies and stores three-dimensional models, panoramas, and interactive photo data according to the dynasties of ancient paintings (such as the Tang Dynasty and Song Dynasty), themes (such as landscape paintings and figure paintings), and author dimensions. At the same time, the data interface is open (for the exhibition hall's online exhibition hall to call, so that visitors who are not present can also view the three-dimensional models of ancient paintings).
[0026] To sum up, in calligraphy and painting exhibition halls, the system uses the binocular vision image acquisition module to obtain multi-perspective images of ancient paintings; the AI image stitching and fusion module processes them to generate panoramic images, and the cultural relics three-dimensional reconstruction module constructs a three-dimensional model that fits the original object; the real-time fusion smart shooting module allows visitors to interact naturally with the ancient painting models, and the interactive output and management module supports results sharing and data retrieval; the system not only provides data support for the protection of ancient paintings, but also enhances the cultural communication effect in an interactive form.
[0027] Example 2: Interactive photo-taking scene in the Bronze Exhibition Hall.
[0028] In a museum's bronze exhibition hall, staff deployed a binocular vision-based AI smart photography system to allow visitors to interact with and take photos of a precious bronze tripod. The specific operation process is as follows: Binocular vision image acquisition module: Based on the layout of the bronze tripod display case and the shape characteristics of the bronze tripod, the staff installed and debugged binocular cameras at appropriate locations around the display case. The system presets a timing trigger strategy and an induction trigger strategy. The timing trigger strategy starts the camera shooting at a fixed interval of 5 minutes (the interval can be adjusted within 1-10 minutes through the exhibition hall management terminal, and the update is completed within 30 seconds after adjustment); at the same time, an array infrared sensor is installed within 1-3 meters in front of the display case. When a visitor stays in front of the display case for more than 5 seconds, the induction trigger strategy will immediately trigger the binocular camera. The camera shoots the bronze tripod from multiple directions and angles, and obtains and stores high-definition original images from different perspectives.
[0029] AI image stitching and fusion module: After receiving these multi-view images, Figure 2 As shown in the figure, denoising and color correction preprocessing are performed first to remove noise in the image and unify the color style. Then, multi-scale feature flow matching is used to align the images. The calculation formula of the multi-scale feature flow matching algorithm is: , so that images from different perspectives are aligned in space; then the registered images are fused with the help of the depth perception fusion function, and the calculation formula of the depth perception fusion function is: , and finally generated a complete and clear panoramic view of the bronze tripod.
[0030] Cultural Relics 3D Reconstruction Module: The three-dimensional grid model is constructed using the image data obtained by the binocular vision image acquisition module. First, the parallax between images from different perspectives is calculated based on the binocular stereo vision principle. The calculation formula is: , the parallax generated by multi-angle shooting is dynamically adjusted through dynamic parallax compensation related calculations; the depth information of the bronze tripod surface points is obtained by combining the triangulation principle and the adjusted parallax to construct point cloud data; then, the curvature-guided triangulated model is combined to generate triangular facets. The calculation formula of the curvature-guided triangulated model is: , screening probability The vertex group is used to generate triangular faces. After adjusting the mesh accuracy to form a three-dimensional mesh model that adapts to the multi-scale characteristics of the bronze tripod, the panoramic image generated by the AI image stitching and fusion module is introduced to extract the structural features of the surface vertices of the three-dimensional mesh model to determine the curvature properties. The panoramic image is texture segmented and the texture fragments are mapped to the corresponding areas of the model. The texture stretching is reduced in the ears and decorative areas with large curvature, and the texture is guaranteed to be complete in the body area of the tripod with small curvature. Finally, a three-dimensional model of the bronze tripod with fine texture is generated.
[0031] Real-time fusion smart shooting module: When a visitor wants to interact with the bronze tripod and take a photo, the real-time fusion intelligent shooting module retrieves the bronze tripod 3D model. At the same time, the camera collects the visitor's real-time image and pre-processes it. The system then fuses the visitor's image with the bronze tripod 3D model using the dynamic interactive anchor fusion algorithm. The calculation formula of the dynamic interactive anchor fusion algorithm is: , and can recognize visitors’ waving and heart-shaped gestures, generating interactive photo screens such as visitors “touching” the ears of the tripod and “standing in the same frame” with the tripod.
[0032] Interactive output and management module: After receiving the photo results, a preview function is provided on the display screen next to it. Visitors can choose to save the photos to their mobile phones or share them on social platforms. At the same time, the system stores relevant data according to the age and type of the bronze tripod, and opens an interface for subsequent digital display and research calls by the exhibition hall.
[0033] To sum up, the AI smart shooting system of the cultural and museum exhibition hall based on binocular vision is used in the bronze exhibition hall to obtain multi-perspective high-definition images through the binocular vision image acquisition module, laying the foundation for subsequent processing. The AI image stitching and fusion module generates panoramic images, and the cultural relics 3D reconstruction module constructs a 3D model with fine textures. The real-time fusion smart shooting module realizes the interactive fusion of visitors and cultural relics models, and the interactive output and management module provides convenient operation and data management. The system not only realizes the digital recording of cultural relics, but also enhances the exhibition experience through interaction.
[0034] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as above in terms of a preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can, without departing from the scope of the technical solution of the present invention, make some changes or modifications to equivalent embodiments using the technical contents disclosed above. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.
Claims
1. The AI smart photography system for cultural and museum exhibition halls based on binocular vision is characterized by: The system includes: binocular vision image acquisition module, AI image stitching and fusion module, cultural relics 3D reconstruction module, real-time fusion intelligent shooting module and interactive output and management module; The binocular vision image acquisition module uses a binocular camera, which is installed and debugged according to the layout of the cultural relics display cabinet and the characteristics of the cultural relics. It shoots cultural relics from multiple directions and angles according to preset strategies, and obtains and stores high-definition original images of cultural relics from different perspectives; The AI image stitching and fusion module receives multi-view high-definition images collected by binoculars and performs denoising and color correction preprocessing; registers the images using a multi-scale feature flow matching algorithm, and fuses the registered images with the help of a depth perception fusion function to generate a panoramic image of the cultural relics; The cultural relic 3D reconstruction module: builds a 3D grid model based on the image data of the binocular vision image acquisition module, introduces the panoramic image of the AI image stitching and fusion module, and generates a 3D cultural relic model with fine texture through texture and structure symbiotic optimization; The real-time fusion intelligent shooting module retrieves the 3D model of the cultural relics 3D reconstruction module, collects and pre-processes the real-time image of the visitor, fuses the two through a dynamic interactive anchor point fusion algorithm, and simultaneously recognizes the visitor's movements to generate an interactive photo screen; The interactive output and management module receives the photo results of the real-time fusion smart shooting module, provides preview, save and sharing functions, stores cultural relics data by dimension classification and opens interfaces for calling.
2. The binocular vision-based AI smart photography system for cultural and museum exhibition halls according to claim 1 is characterized in that: In the binocular vision image acquisition module, preset strategies include a timing trigger strategy and an induction trigger strategy; The timing trigger strategy is: start the binocular camera shooting at a fixed time interval, and the time interval is dynamically adjusted within the range of 1 minute to 10 minutes through the exhibition hall management terminal. After the adjustment, the system completes the strategy update within 30 seconds; The induction triggering strategy is: using an array of infrared sensors to detect the human body within a range of 1 to 3 meters in front of the cultural relics display cabinet, and triggering the binocular camera to shoot when it is recognized that the human body stays for ≥5 seconds.
3. The binocular vision-based AI smart photography system for cultural and museum exhibition halls according to claim 1 is characterized in that: In the AI image stitching and fusion module, the multi-scale feature flow matching algorithm is used to align the images. The calculation formula is: ,in are the left and right view feature points, is the scale weight, Control the degree of penalty for flow field differences, and Both are left and right perspectives The feature vector of the scale, is the L2 norm.
4. The binocular vision-based AI smart photography system for cultural and museum exhibition halls according to claim 1 is characterized in that: In the AI image stitching and fusion module, the registered images are fused with the help of the depth perception fusion function. The calculation formula is: ,in The fused image is The pixel value at and Both are images after left and right view registration. is the fusion weight, and the calculation formula is: , and Both left and right perspective images are in The depth value at is the depth difference threshold, is the hyperbolic tangent function.
5. The binocular vision-based AI smart photography system for cultural and museum exhibition halls according to claim 1 is characterized in that: In the cultural relic 3D reconstruction module, the specific steps of constructing a 3D grid model based on the image data of the binocular vision image acquisition module are as follows: based on the image data acquired by the binocular vision image acquisition module, the principle of binocular stereo vision is used to calculate the parallax between images of different viewing angles; through dynamic parallax compensation related calculations, the parallax generated by multi-angle shooting is dynamically adjusted; then, based on the principle of triangulation, the depth information of the surface points of the cultural relic is obtained in combination with the adjusted parallax, and the point cloud data of the cultural relic is constructed; the curvature is used to guide the triangulated model to generate triangular facets, and the grid accuracy is adjusted to form a 3D grid model that adapts to the multi-scale characteristics of the cultural relic.
6. The binocular vision-based AI smart photography system for cultural and museum exhibition halls according to claim 5 is characterized in that: In the cultural relic 3D reconstruction module, the calculation formula for the parallax between images from different perspectives after compensation is: ,in, It is Collection time After compensation, the parallax at is the original disparity, For the The viewing angle offset of the acquisition, and is the compensation coefficient, which controls the effect of the offset on the compensation result. is the gradient of the original disparity.
7. The binocular vision-based AI smart photography system for cultural and museum exhibition halls according to claim 5 is characterized in that: In the cultural relic 3D reconstruction module, the curvature-guided triangulated model is used to generate triangular facets. The calculation formula is: ,in It is a three-point group The triangulated probability of is the curvature value corresponding to the three points, and the probability of triangulation of the three-point group is calculated. is the curvature penalty coefficient, is the Euclidean distance, the screening probability The vertex group generates triangles.
8. The binocular vision-based AI smart photography system for cultural and museum exhibition halls according to claim 1 is characterized in that: In the cultural relic 3D reconstruction module, a 3D model of a cultural relic with fine texture is generated through texture structure symbiosis optimization. The specific steps are as follows: (1) Obtain the panoramic image of cultural relics generated by the AI image stitching and fusion module as the original material for texture mapping; (2) Extract the structural features corresponding to each vertex on the surface of the 3D mesh model and determine the curvature properties of different areas on the model surface; (3) Perform texture segmentation on the panoramic image to divide the texture segments corresponding to the different curvature areas of the 3D mesh model; (4) Map the segmented texture segments to the corresponding areas of the 3D mesh model so that the texture fits the surface structure of the model; (5) Adjust the fit of the texture on the model surface, reduce texture stretching in areas where the curvature value K>0.8, and ensure texture integrity in areas where the curvature value K≤0.8; (6) After completing the texture mapping of all areas, integrate them to form a three-dimensional model of the cultural relic with fine textures.
9. The binocular vision-based AI smart photography system for cultural and museum exhibition halls according to claim 1 is characterized in that: In the real-time fusion smart shooting module, the calculation formula of the dynamic interaction anchor point fusion algorithm is: ,in, is the real-time coordinate of the visitor's hand key point in the camera coordinate system, The coordinates of the interaction anchor point of the cultural relic 3D model are, It is the coordinates of the real-time coordinates of the visitor's hand key points in the camera coordinate system and the coordinates of the interactive anchor points of the cultural relic 3D model. is the Euclidean distance between the hand and the initial position of the anchor point, The default position of the anchor point in a static state is the fusion sharpness parameter, is the dynamic scaling factor, is the maximum interaction distance threshold.
Citation Information
Patent Citations
Gaze point estimation method and system, storage medium and terminal
CN109407828A
Binocular parking identification method and system based on image stitching, and processor
CN118015599A
Artificial intelligence assisted cultural relic digital reproduction system
CN119027619A
Improved binocular stereo matching fusion algorithm
CN119251198A
Cultural relic handicraft three-dimensional reconstruction method based on multiple views and deep learning
CN120451444A
Cited By
Pipeline construction monitoring method and system based on image recognition
CN121982000A
Pipeline construction monitoring method and system based on image recognition
CN121982000B