AI intelligent shooting system for cultural and educational exhibition hall based on binocular vision
The AI-powered smart photography system for cultural heritage exhibition halls, based on binocular vision, integrates multimodal technology to generate high-precision 3D models of cultural relics and supports real-time interaction. This solves the problems of limited perspective and insufficient interaction in traditional cultural heritage exhibition halls, thereby improving the visitor experience and resource management efficiency.
Patent Information
- Application Number
- CN202511140820.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-15
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-08-15
AI Technical Summary
Traditional museum and cultural exhibition technologies suffer from problems such as limited perspectives, incomplete information coverage, high costs, low efficiency, and difficulty in supporting real-time interaction and cross-platform sharing, resulting in low audience participation and difficulties in resource management.
The museum adopts an AI-powered intelligent photography system based on binocular vision, which includes a binocular vision image acquisition module, an AI image stitching and fusion module, a cultural relic 3D reconstruction module, and a real-time fusion intelligent photography module. It generates high-precision panoramic images and 3D models through multi-strategy triggered camera shooting, multi-scale feature flow matching, and depth perception fusion algorithms, and enables visitors to take interactive photos with cultural relics.
It has achieved an intelligent upgrade in the display of cultural relics and interaction with visitors, generated high-precision 3D models of cultural relics, enhanced visitors' immersion and participation, and supported the classified storage and cross-platform sharing of data, thereby improving the digital operation and cultural dissemination effect of the exhibition hall.
Smart Images

Figure CN120726575B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of digitalization and interaction of cultural relics, in particular to an AI intelligent shooting system for cultural relics exhibition based on binocular vision. BACKGROUND
[0002] With the vigorous development of cultural tourism industry, as an important carrier for the protection and dissemination of cultural heritage, the display means and audience interactive experience of cultural relics exhibition have been increasingly improved. Traditional exhibition methods rely on static objects and text descriptions, which are difficult to fully present the details and historical background of cultural relics, especially in the three-dimensional appearance and texture characteristics. At the same time, the breakthroughs in computer vision, artificial intelligence and three-dimensional reconstruction technology provide new solutions for the digitalization of cultural relics. Based on the binocular vision stereoscopic imaging technology, multi-angle high-precision image acquisition can be realized by simulating the principle of human eye parallax. Combined with deep learning algorithm image fusion and three-dimensional modeling technology, a realistic three-dimensional model can be further generated. Therefore, how to integrate multi-modal technology into cultural relics exhibition has become a key issue to improve the exhibition effect and audience participation.
[0003] The shooting and display technology of traditional cultural relics exhibition mainly relies on monocular camera or fixed-angle photography equipment, which has the problems of single view angle and incomplete information coverage, and it is difficult to fully present the complex structure and detail texture of cultural relics. For example, the planar image shot by monocular camera cannot reflect the three-dimensional appearance of cultural relics, and artificial three-dimensional reconstruction needs to rely on professional equipment and complex operation process, which is high in cost and low in efficiency. In addition, traditional technology usually cannot support real-time interaction function, which limits the audience participation and makes it difficult to generate personalized content through motion interaction, resulting in serious homogeneity of visit experience. In terms of resource management, traditional systems mostly use decentralized storage, which has single dimension of cultural relic data and is difficult to call, and cannot meet the needs of cross-platform sharing and secondary development. These limitations restrict the digital transformation and innovative development of cultural relics exhibition. SUMMARY
[0004] The purpose of the present application is to overcome the shortcomings of the prior art and provide an AI intelligent shooting system for cultural relics exhibition based on binocular vision, which includes a binocular vision image acquisition module that can shoot cultural relics from multiple directions; an AI image stitching and fusion module that can denoise, correct and register and fuse images to generate a panoramic image; a cultural relic three-dimensional reconstruction module that constructs a three-dimensional grid model and generates a three-dimensional model with fine texture after optimization; a real-time fusion intelligent shooting module that fuses three-dimensional models and visitor images to generate interactive images; and an interactive output and management module that provides preview function and stores data in categories.
[0005] The application provides the following technical solutions to solve the above technical problems: an AI intelligent photographing system for a cultural exhibition hall based on binocular vision, which comprises a binocular vision image acquisition module, an AI image splicing and fusion module, a cultural relic three-dimensional reconstruction module, a real-time fusion intelligent photographing module, and an interactive output and management module.
[0006] The binocular vision image acquisition module: uses a binocular camera, installs and adjusts according to the layout of cultural relic display cabinets and the characteristics of cultural relics, and photographs cultural relics from multiple directions and angles according to a preset strategy to obtain and store high-definition original images of cultural relics from different perspectives.
[0007] The AI image splicing and fusion module: receives multi-perspective high-definition images collected by the binocular camera, performs denoising and color correction preprocessing, uses a multi-scale feature flow matching algorithm to register the images, fuses the registered images by means of a depth perception fusion function, and generates a cultural relic panorama.
[0008] The cultural relic three-dimensional reconstruction module: constructs a three-dimensional grid model based on the image data of the binocular vision image acquisition module, introduces the panorama image of the AI image splicing and fusion module, and generates a cultural relic three-dimensional model with fine texture through texture and structure symbiotic optimization.
[0009] The real-time fusion intelligent photographing module: calls the three-dimensional model of the cultural relic three-dimensional reconstruction module, collects and preprocesses real-time images of visitors, fuses the two through a dynamic interactive anchor point fusion algorithm, and simultaneously identifies the actions of the visitors to generate an interactive photographing picture.
[0010] The interactive output and management module: receives the photographing results of the real-time fusion intelligent photographing module, provides preview, saving, and sharing functions, stores cultural relic data by dimension and opens an interface for calling.
[0011] Further, in the binocular vision image acquisition module, the preset strategy comprises a timing trigger strategy and a sensing trigger strategy.
[0012] The timing trigger strategy: starts the binocular camera to photograph at fixed time intervals, the time interval can be dynamically adjusted by the exhibition hall management terminal within the range of 1 minute to 10 minutes, and the system completes strategy updating within 30 seconds after adjustment.
[0013] The sensing trigger strategy: detects the human body staying state within 1-3 meters in front of the cultural relic display cabinet through an arrayed infrared sensor, and triggers the binocular camera to photograph when the identified human body staying time is greater than or equal to 5 seconds.
[0014] Further, in the AI image splicing and fusion module, the multi-scale feature flow matching algorithm is used to register the images, and the calculation formula is: wherein is a left-view perspective feature point, is a scale weight, a penalty of a control flow field difference, and are feature vectors of the first scale of the left and right view angles, is an L2 norm.
[0015] Further, in the AI image stitching and fusion module, the registered images are fused by means of a depth perception fusion function, and the calculation formula is: wherein is the pixel value of the fused image at and are the registered images of the left and right view angles, is a fusion weight, and the calculation formula is: and are the depth values of the left and right view angle images at is a depth difference threshold, is a hyperbolic tangent function.
[0016] Further, in the cultural relic three-dimensional reconstruction module, the specific steps of constructing a three-dimensional mesh model based on the image data of the binocular vision image acquisition module are as follows: according to the image data obtained by the binocular vision image acquisition module, the parallax between different view angle images is calculated by using the binocular stereo vision principle, the parallax generated by multi-angle shooting is dynamically adjusted through dynamic parallax compensation related calculation, and then the depth information of the cultural relic surface points is obtained according to the triangulation principle combined with the adjusted parallax, the point cloud data of the cultural relic is constructed, the triangular facets are generated by combining the curvature guided triangulation model, and the mesh precision is adjusted to form a three-dimensional mesh model that adapts to the multi-scale characteristics of the cultural relic.
[0017] Further, in the cultural relic three-dimensional reconstruction module, the calculation formula of the compensated parallax between different view angle images is: wherein is the compensated parallax at is the original parallax, is the view angle offset of the th acquisition, and is a compensation coefficient that controls the influence of the offset on the compensation result, is the gradient of the original parallax.
[0018] Further, in the cultural relic three-dimensional reconstruction module, the triangular facets are generated by combining the curvature guided triangulation model, and the calculation formula is: wherein is a three-point group triangulation probability of a triangle, is the curvature value of a three-point correspondence, and the triangulation probability of a three-point group is calculated, is a curvature penalty coefficient, is the Euclidean distance, and the screening probability generates a triangular facet from the vertex group.
[0019] Further, in the cultural relic three-dimensional reconstruction module, the fine-textured cultural relic three-dimensional model is generated through texture structure coexistence optimization, and the specific steps are as follows:
[0020] (1) Obtain the cultural relic panoramic image generated by the AI image stitching and fusion module as the original material for texture mapping;
[0021] (2) Extract the structural features corresponding to each vertex on the surface of the three-dimensional mesh model, and determine the curvature properties of different regions on the model surface;
[0022] (3) Texture segmentation is performed on the panoramic image to divide texture segments corresponding to different curvature regions of the three-dimensional mesh model;
[0023] (4) Map the segmented texture segments to the corresponding regions of the three-dimensional mesh model to make the texture fit the surface structure of the model;
[0024] (5) Adjust the texture fitting state on the model surface, reduce texture stretching in regions with curvature value K>0.8, and ensure texture integrity in regions with curvature value K≤0.8, wherein the curvature value is calculated as follows: select the 1-ring neighborhood vertices of the triangulation mesh vertex of the cultural relic three-dimensional model to form a local region; by analyzing the surface morphology of the local region, a quantitative index reflecting the surface bending feature is obtained as the curvature value of the vertex, which is used to measure the bending degree of the model surface; the curvature threshold is determined as follows: select a test model constructed from cultural relics and exhibits, perform three-dimensional reconstruction by the system, and analyze the texture stretching deformation under different curvature values; when the curvature value exceeds a certain critical value, the texture stretching rate increases significantly, and the critical value is defined as the curvature threshold; it has been verified that when the curvature value is greater than 0.8, the texture stretching risk meets the requirement of reducing stretching, and when the curvature value is less than or equal to 0.8, the texture fitting state meets the requirement of ensuring integrity, so 0.8 is used to distinguish regions with large curvature and regions with small curvature;
[0025] (6) After completing the texture mapping of all regions, a fine-textured cultural relic three-dimensional model is formed.
[0026] Further, in the real-time fusion intelligent shooting module, the calculation formula of the dynamic interactive anchor point fusion algorithm is: wherein, is the real-time coordinate of the visitor's hand key point in the camera coordinate system, is the interactive anchor point coordinate of the cultural relic three-dimensional model, is the real-time coordinate of the visitor's hand key point in the camera coordinate system and the interactive anchor point coordinate of the cultural relic three-dimensional model after fusion, is the Euclidean distance between the hand and the initial position of the anchor point, is the default position of the anchor point in the static state is the fusion sharpness parameter, is the dynamic scaling coefficient, is the maximum interactive distance threshold.
[0027] Compared with the prior art, the AI intelligent shooting system for cultural relics exhibition based on binocular vision has the following beneficial effects:
[0028] Firstly, the AI intelligent shooting system for cultural relics exhibition based on binocular vision integrates binocular vision image acquisition, AI image splicing and fusion, cultural relic three-dimensional reconstruction and real-time fusion intelligent shooting technology modules, realizes the intelligent upgrading of cultural relic display and visitor interaction, adopts a multi-strategy triggered binocular camera shooting mechanism, combines the timing and sensing trigger mode, can flexibly adapt to the needs of different exhibition scenes, ensures the efficiency of cultural relic image acquisition, and avoids resource waste, through multi-scale feature flow matching and depth perception fusion algorithm, the system can generate high-precision, low-distortion cultural relic panorama and three-dimensional model, significantly improves the problem of single view angle and easy loss of details in traditional shooting, and provides more rich data support for cultural relic digital protection and research.
[0029] Secondly, through the real-time fusion intelligent shooting module, the system can collect visitor images and dynamically fuse with the three-dimensional model of cultural relics, generate personalized interactive shooting pictures combined with motion recognition technology, break through the limitations of traditional static display, enhance the immersion and participation of visitors, at the same time, the interactive output and management module provides one-stop services such as preview, save and share, and supports cultural relic data classification storage and interface calling, which provides a convenient solution for digital operation and cultural dissemination of the exhibition hall.
[0030] Other advantages, objects and features of the present application will be set forth in part in the description which follows, and in part will become apparent to those skilled in the art upon examination of the following or can be learned by practice of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can also obtain other drawings according to these drawings without creative labor.
[0032] Figure 1A flowchart of an AI intelligent shooting system for a cultural exhibition hall based on binocular vision;
[0033] Figure 2 A data flow conversion framework diagram of an AI intelligent shooting system for a cultural exhibition hall based on binocular vision. DETAILED DESCRIPTION
[0034] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined inventive purpose, the specific embodiments, structures, features and effects thereof according to the present application will be described in detail below in combination with the drawings and preferred embodiments.
[0035] Embodiment one:
[0036] An interactive experience scene of a calligraphy and painting exhibition hall.
[0037] A calligraphy and painting exhibition hall introduces an AI intelligent shooting system for a cultural exhibition hall based on binocular vision, aiming to enable visitors to have immersive interaction with ancient paintings through digital technology, and at the same time, to digitally record and protect the ancient paintings. The specific process is as follows:
[0038] Binocular vision image acquisition module:
[0039] The staff installs and adjusts the binocular camera at a position 1.5 meters in front of the exhibition cabinet and at a 45° angle position on both sides according to the hanging height of the ancient painting exhibition cabinet, the size of the ancient painting and the distribution of the picture details. The system has preset a timing trigger strategy and a sensing trigger strategy: the timing trigger strategy automatically takes pictures at an interval of 3 minutes (the interval can be adjusted by the management terminal within 1-10 minutes, and the adjustment takes effect within 30 seconds), ensuring that the state of the ancient painting at different time periods can be recorded; an array type infrared sensor is deployed within a range of 1-3 meters in front of the exhibition cabinet, and when a visitor stays in front of the exhibition cabinet for ≥5 seconds (indicating that he / she has strong interest in the ancient painting), the sensing trigger strategy triggers the camera to take pictures immediately. The binocular camera takes pictures from multiple directions and angles, and finally obtains and stores 20 groups of high-definition original images of the ancient painting from different perspectives, as shown in FIG. Figure 1
[0040] AI image stitching and fusion module
[0041] The multi-view images are received, and first, denoising processing (to remove light spots and camera sensor noise caused by light reflection during shooting) and color correction (to unify the brightness and color temperature of different view images, and restore the ink color density and color level of the ancient painting) are performed. Then, the images are registered using multi-scale feature flow matching. The calculation formula of the multi-scale feature flow matching algorithm is as follows: wherein is the left and right view feature points, is the scale weight, is the punishment degree of flow field difference control, and are the left and right view feature points. a feature vector of a scale, is the L2 norm, by identifying the edge lines of ancient paintings, seal profile features, and accurately aligning different perspective images in spatial position (such as ensuring that the left perspective frame edge seamlessly connects with the front perspective frame edge); then the registered images are fused by means of a depth perception fusion function, and the calculation formula of the depth perception fusion function is: wherein is the pixel value of the fused image at , and are the left and right perspective registered images, is the fusion weight, and the calculation formula is: , and are the depth values of the left and right perspective images at , is the depth difference threshold, is the hyperbolic tangent function, and finally a complete ancient painting panorama is generated, and the traces of inscriptions and the details of the brush strokes of the landscape are clearly visible.
[0042] Cultural relic three-dimensional reconstruction module:
[0043] First, a three-dimensional grid model is constructed using the image data of the binocular vision image acquisition module, and the parallax between different perspective images is calculated according to the binocular stereo vision principle, and the calculation formula is: wherein, is the compensated parallax at when the th acquisition is performed, is the original parallax, is the perspective offset amount of the th acquisition, and are compensation coefficients that control the influence of the offset amount on the compensation result, is the gradient of the original parallax, and the parallax deviation caused by the difference in shooting angle is eliminated by dynamic parallax compensation related calculations; the depth information of the surface points of the ancient painting (including the frame) is obtained by combining the triangulation principle, and the point cloud data is constructed; then a curvature-guided triangulation model is used to generate triangular facets (select vertex groups that meet the triangulation probability standard), and the calculation formula of the curvature-guided triangulation model is: wherein is the triangulation probability of the three-point group , is the curvature value corresponding to the three points, and the triangulation probability of the three-point group is calculated, is the curvature penalty coefficient, is the Euclidean distance, and the probability The vertex group generates a triangular facet, adjusts the grid accuracy, forms a three-dimensional grid model that can reflect the three-dimensional sense of the picture frame and the flatness of the paper, and then introduces the panoramic image of the AI image splicing fusion module. The structural features (such as the corners of the picture frame and the curvature of the paper creases) of the surface vertices of the three-dimensional grid model are extracted, the panoramic image is texture segmented (the picture frame texture and the picture main body texture are separated), and the texture fragments are correspondingly mapped to the corresponding areas of the model. In the area with large curvature of the picture frame corner, the texture stretching is reduced (to avoid deformation of the picture frame wood grain), and in the area with small curvature of the picture main body, the texture is ensured to be complete (to ensure the coherence of the landscape pattern). Finally, a three-dimensional model of an ancient painting with fine texture is generated.
[0044] Real-time fusion intelligent shooting module:
[0045] When the visitor stands in the designated interactive area of the exhibition hall (a camera is provided in front), the real-time fusion intelligent shooting module collects the real-time image of the visitor through the camera (and performs preprocessing such as removing background color and adjusting the brightness of the portrait), the system calls the three-dimensional reconstruction module of the cultural relics to generate a three-dimensional model of an ancient painting, and fuses the image of the visitor with the three-dimensional model of the ancient painting through a dynamic interactive anchor point fusion algorithm. The calculation formula of the dynamic interactive anchor point fusion algorithm is: wherein, is the real-time coordinate of the visitor's hand key point in the camera coordinate system, is the interactive anchor point coordinate of the three-dimensional model of cultural relics, is the coordinate after fusion of the real-time coordinate of the visitor's hand key point in the camera coordinate system and the interactive anchor point coordinate of the three-dimensional model of cultural relics, is the Euclidean distance between the hand and the initial position of the anchor point, is the default position of the anchor point in the static state is the fusion sharpness parameter, is the dynamic scaling coefficient, is the maximum interactive distance threshold value, for example, when the visitor stretches out his hand, the system recognizes the hand action, fuses the image of the hand with the interactive anchor point of the "scroll" in the ancient painting, and forms the visual effect of "holding the scroll"; if the visitor makes a bowing gesture, the system will generate an interactive picture of "bowing to the ancient people in the picture".
[0046] Interactive output and management module:
[0047] After receiving the shooting results of the real-time fusion intelligent shooting module, the preview is provided on the display screen in the interactive area (for visitors to view the shooting effect), while supporting code scanning and saving (saving the photos to the mobile phone album) and social platform sharing (generating a sharing link with the exhibition hall logo), the system classifies and stores the three-dimensional model, panoramic map and interactive photo data according to the dynasty (such as the Tang Dynasty and the Song Dynasty), theme (such as landscape painting and figure painting) and author dimensions of ancient paintings, and opens a data interface (for online exhibition halls to call, so that visitors who have not arrived on site can also view the three-dimensional model of ancient paintings).
[0048] In summary, in the calligraphy and painting exhibition hall, the system acquires multi-view images of ancient paintings by means of the binocular vision image acquisition module; generates a panoramic map by processing through the AI image stitching and fusion module, and constructs a three-dimensional model that fits the original object by the cultural relic three-dimensional reconstruction module; the real-time fusion intelligent shooting module allows visitors to interact naturally with the ancient painting model, and the interactive output and management module supports result sharing and data calling; the system not only provides data support for the protection of ancient paintings, but also enhances the cultural communication effect in an interactive form.
[0049] Embodiment Two:
[0050] Bronze exhibition hall interactive shooting scene.
[0051] In the bronze exhibition hall of a certain museum, the staff deployed an AI intelligent shooting system based on binocular vision for the visitors to have immersive interactive shooting with the precious bronze ding, and the specific operation process is as follows:
[0052] Binocular vision image acquisition module:
[0053] The staff installed and debugged the binocular camera at appropriate positions around the bronze ding showcase according to the layout of the bronze ding showcase and the modeling features of the bronze ding, the system preset a timing trigger strategy and a sensing trigger strategy, the timing trigger strategy started the camera shooting at a fixed interval of 5 minutes (the interval can be adjusted by the exhibition hall management terminal within 1-10 minutes, and the update is completed within 30 seconds after adjustment); at the same time, an array type infrared sensor was installed 1-3 meters away from the front of the showcase, when a visitor stayed in front of the showcase for more than 5 seconds, the sensing trigger strategy would trigger the binocular camera immediately, the camera shot the bronze ding from multiple directions and angles, and acquired and stored high-definition original images of different perspectives.
[0054] AI image stitching and fusion module:
[0055] After receiving these multi-view images, as shown in Figure 2 , first, the noise in the image is removed and the color style is unified, then the images are registered by using multi-scale feature flow matching, and the calculation formula of the multi-scale feature flow matching algorithm is: aligning the images of different perspectives in spatial position; then fusing the registered images by means of a depth perception fusion function, and the calculation formula of the depth perception fusion function is: , and finally generating a complete and clear panoramic view of the bronze ding.
[0056] The cultural relic three-dimensional reconstruction module:
[0057] The image data obtained by the binocular vision image acquisition module is used to construct a three-dimensional grid model. First, the parallax between images of different perspectives is calculated according to the binocular stereo vision principle, and the calculation formula is: The parallax generated by multi-angle shooting is dynamically adjusted through dynamic parallax compensation related calculation; then the depth information of the bronze ding surface points is obtained by combining the triangulation principle and the adjusted parallax to construct point cloud data; then the triangular facets are generated by combining the curvature guided triangulation model, and the calculation formula of the curvature guided triangulation model is: , and the triangular facets are generated by screening the vertex groups with a probability After adjusting the grid precision to form a three-dimensional grid model that adapts to the multi-scale features of the bronze ding, the panoramic image generated by the AI image stitching and fusion module is introduced, the structural features of the surface vertices of the three-dimensional grid model are extracted to determine the curvature properties, the texture segmentation is performed on the panoramic image, and the texture segments are correspondingly mapped to the corresponding areas of the model, the texture stretching is reduced in the areas with large curvature such as the ding ears and decorative patterns, and the texture is ensured to be complete in the areas with small curvature such as the ding body, and finally the bronze ding three-dimensional model with fine texture is generated.
[0058] Real-time fusion intelligent shooting module:
[0059] When the visitor wants to interact with the bronze ding and take a photo, the real-time fusion intelligent shooting module calls the bronze ding three-dimensional model, and at the same time, the real-time image of the visitor is collected by the camera and preprocessed, the system fuses the image of the visitor with the bronze ding three-dimensional model through a dynamic interactive anchor point fusion algorithm, and the calculation formula of the dynamic interactive anchor point fusion algorithm is: , and can recognize the visitor's hand waving and heart touching actions, and generate interactive photo pictures such as the visitor "touching" the ding ears and taking photos with the ding.
[0060] Interactive output and management module:
[0061] After receiving the shooting results, the preview function is provided on the display screen next to it, the visitor can choose to save the photos to the mobile phone or share them on social platforms, at the same time, the system stores the related data according to the age and type dimensions of the bronze ding, and opens the interface for subsequent digital display and research calling of the museum.
[0062] In summary, the AI intelligent shooting system of the cultural relics museum based on binocular vision acquires multi-view high-definition images through the binocular vision image acquisition module in the bronze ware museum, lays a foundation for subsequent processing, generates panoramic images through the AI image splicing and fusion module, constructs three-dimensional models with fine textures through the cultural relic three-dimensional reconstruction module, realizes the interactive fusion of visitors and cultural relic models through the real-time fusion intelligent shooting module, and provides convenient operation and data management through the interactive output and management module. The system not only realizes the digital recording of cultural relics, but also improves the exhibition experience through interaction.
[0063] The above is only a preferred embodiment of the present application, and does not limit the present application in any form. Although the present application has been disclosed as above with a preferred embodiment, it is not intended to limit the present application. Any person skilled in the art can make slight changes or modifications to the above disclosed technical content to obtain equivalent embodiments with equivalent changes, without departing from the technical solution of the present application. Any modification, equivalent change and modification of the above embodiments made in accordance with the technical essence of the present application still fall within the scope of the technical solution of the present application.
Claims
1. A binocular vision-based AI intelligent shooting system for a cultural and historical exhibition hall, characterized in that, The system comprises a binocular vision image acquisition module, an AI image splicing and fusion module, a cultural relic three-dimensional reconstruction module, a real-time fusion intelligent shooting module and an interactive output and management module. The binocular vision image acquisition module: uses a binocular camera, installs and adjusts according to the layout of a cultural relic showcase and the characteristics of cultural relics, shoots cultural relics from multiple directions and angles according to a preset strategy, and acquires and stores high-definition original images of cultural relics from different angles. The AI image splicing and fusion module: receives multi-angle high-definition images collected by the binocular camera, carries out denoising and color correction preprocessing, uses a multi-scale feature flow matching algorithm to register the images, fuses the registered images by means of a depth perception fusion function, and generates a cultural relic panorama. The cultural relic three-dimensional reconstruction module: constructs a three-dimensional grid model based on the image data of the binocular vision image acquisition module, introduces the panorama image of the AI image splicing and fusion module, and generates a cultural relic three-dimensional model with fine texture through texture and structure symbiotic optimization. The specific steps for constructing a 3D mesh model based on image data from a binocular vision image acquisition module are as follows: Using the image data acquired by the binocular vision image acquisition module and applying the principle of binocular stereo vision, calculate the disparity between images from different viewpoints. The calculation formula is as follows: ,in, It is the first During the next collection Parallax after compensation at the location, For the original parallax, For the first The viewpoint offset of the second acquisition. and It is a compensation coefficient that controls the impact of the offset on the compensation result. This is the gradient of the original parallax. Through dynamic parallax compensation calculations, the parallax generated by multi-angle shooting is dynamically adjusted. Then, based on the principle of triangulation, the depth information of the surface points of the artifact is obtained by combining the adjusted parallax, constructing the point cloud data of the artifact. Combined with the curvature-guided triangulation model, triangular patches are generated. The calculation formula is: ,in It is a three-point group The triangulation probability, Given the curvature values corresponding to the three points, calculate the triangulation probability of the three-point group. It is the curvature penalty coefficient. It's Euclidean distance, the selection probability. The vertex groups generate triangular patches and the mesh precision is adjusted to form a three-dimensional mesh model adapted to the multi-scale features of cultural relics. The real-time fusion intelligent shooting module: calls the three-dimensional model of the cultural relic three-dimensional reconstruction module, collects and preprocesses real-time images of visitors, fuses the two through a dynamic interactive anchor point fusion algorithm, simultaneously identifies the actions of the visitors, and generates an interactive shooting picture. The interactive output and management module: receives the shooting results of the real-time fusion intelligent shooting module, provides preview, saving and sharing functions, stores cultural relic data by dimension and opens an interface for calling.
2. The binocular vision-based AI intelligent shooting system for cultural relics and museum according to claim 1, characterized in that, In the binocular vision image acquisition module, the preset strategy includes a timing trigger strategy and a sensing trigger strategy. The timing trigger strategy: starts the binocular camera shooting at fixed time intervals, and the time interval is dynamically adjusted by the exhibition hall management terminal within the range of 1 minute to 10 minutes, and the system completes strategy updating within 30 seconds after adjustment. The sensing trigger strategy: detects the human body staying state within 1-3 meters in front of the cultural relic showcase through an arrayed infrared sensor, and triggers the binocular camera shooting when the identified human body staying time is greater than or equal to 5 seconds.
3. The binocular vision-based AI intelligent shooting system for cultural relics and museum according to claim 1, characterized in that, In the AI image stitching and fusion module, a multi-scale feature flow matching algorithm is used to register the images. The calculation formula is as follows: ,in For left and right view feature points, It is a scale weight. Penalty for controlling flow field differences and Both are left and right perspectives. Feature vectors at each scale It is an L2 norm.
4. The binocular vision-based AI intelligent shooting system for cultural relics and museum according to claim 1, characterized in that, In the AI image stitching and fusion module, the registered images are fused by means of a depth perception fusion function, and the calculation formula is: Wherein is the pixel value of the fused image at , and are the images registered from left and right perspectives, is the fusion weight, and the calculation formula is: , and are the depth values of the left and right perspective images at , is the depth difference threshold, is the hyperbolic tangent function.
5. The binocular vision-based AI intelligent shooting system for cultural relics and museum according to claim 1, characterized in that, In the cultural relic three-dimensional reconstruction module, the three-dimensional model with fine texture of cultural relics is generated through texture structure symbiotic optimization, and the specific steps are as follows: (1) Obtain the cultural relic panorama image generated by the AI image splicing and fusion module as the original material for texture mapping; (2) Extract the structure features corresponding to each vertex on the surface of the three-dimensional grid model, and determine the curvature properties of different regions on the model surface; (3) Perform texture segmentation on the panorama image, and divide the texture segments corresponding to different curvature regions of the three-dimensional grid model; (4) Map the segmented texture segments to the corresponding regions of the three-dimensional grid model to make the texture fit the model surface structure; (5) Adjust the fitting state of the texture on the model surface, reduce the texture stretching in the region with a curvature value K>0.8, and ensure the texture integrity in the region with a curvature value K≤0.8; (6) After completing the texture mapping of all regions, integrate to form a cultural relic three-dimensional model with fine texture.
6. The binocular vision-based AI intelligent shooting system for cultural relics and museum according to claim 1, characterized in that, The calculation formula of the dynamic interaction anchor point fusion algorithm in the real-time fusion intelligent shooting module is: Wherein, is the real-time coordinate of the visitor's hand key point in the camera coordinate system, is the interaction anchor point coordinate of the cultural relic three-dimensional model, is the coordinate after the real-time coordinate of the visitor's hand key point in the camera coordinate system and the interaction anchor point coordinate of the cultural relic three-dimensional model are fused, is the Euclidean distance of the initial position of the hand and the anchor point, is the default position of the anchor point in the static state is the fusion sharpness parameter, is the dynamic scaling coefficient, is the maximum interaction distance threshold.
Citation Information
Patent Citations
Artificial intelligence assisted cultural relic digital reproduction system
CN119027619A
Improved binocular stereo matching fusion algorithm
CN119251198A
Cultural relic handicraft three-dimensional reconstruction method based on multiple views and deep learning
CN120451444A