Scenic area photographing method and device based on artificial intelligence
Through the artificial intelligence-based scenic spot photography method, AI algorithms are used to fuse tourists' facial images into the scenic spot theme template to generate personalized scenic spot fusion images or videos, which solves the problems of insufficient interactivity and cultural integration of traditional photography methods, provides diversified souvenirs, and enhances the tourist experience.
Patent Information
- Application Number
- CN202511015082.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-21
AI Technical Summary
Existing technologies are unable to meet tourists' personalized photography needs at tourist attractions. Traditional photography methods have problems such as poor interactivity, insufficient cultural integration, high costs, and cumbersome processes, and are unable to generate dynamic videos and highly personalized images.
Using an AI-based method, it receives user instructions to capture facial images, selects a theme template, and uses AI algorithms to fuse the facial image into the target theme template to generate a fused image or video of the scenic area. It supports age adjustment and three-dimensional data production, and provides personalized souvenirs.
It achieves highly personalized scenic spot photography results, meets tourists' needs for dynamic videos and in-depth cultural experiences, enhances the interactivity and cultural integration of photography, provides diversified souvenirs, and improves the management order of scenic spots.
Smart Images

Figure CN120825640A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to a scenic spot photographing method and device based on artificial intelligence. Background Art
[0002] Currently, there are two main ways for tourists to take photos at tourist attractions. One is self-portrait, where visitors use their phones or cameras. However, the quality of the photos is limited by their personal equipment and photography skills, and it is difficult to fully integrate the photos with the characteristics of the tourist attraction. The other is traditional photography services, where the scenic spot or a third party provides photographers to take photos. This may require waiting in line, and some commercial photography activities may occupy the best shooting spots for a long time, affecting the experience of other tourists and the order of the scenic spot, and even posing potential risks to the environment or cultural relics. When using these two photography methods, tourists may rent the scenic spot's unique clothing, but the high cost, cumbersome process, and clothing hygiene issues are all troublesome for consumers. Therefore, these technologies are unable to meet tourists' growing demand for personalized photos. Summary of the Invention
[0003] The purpose of the present invention is to provide a scenic spot photography method and device based on artificial intelligence to meet the tourists' growing demand for personalized photography.
[0004] The present invention provides an artificial intelligence-based scenic spot photography method, which includes: receiving a photography instruction sent by a user, and capturing a first facial image of the user according to the photography instruction; receiving a theme template selection instruction sent by the user, and determining a target theme template from a plurality of preset theme templates according to the theme template selection instruction; wherein the target theme template includes at least one of the following: a scenic spot theme, a virtual costume, and a dynamic video template; using a preset artificial intelligence algorithm, the first facial image is fused to the facial area in the target theme template to obtain a first fusion result; based on the first fusion result, a scenic spot photography result of the user is obtained; wherein the scenic spot photography result is: a scenic spot fused image or a scenic spot fused video.
[0005] Furthermore, based on the first fusion result, the step of obtaining the user's scenic spot photo result includes: displaying the first fusion result through a display interface; receiving an age adjustment instruction sent by the user, wherein the age adjustment instruction carries a target age; according to the age adjustment instruction, adjusting the first facial image displayed in the facial area of the first fusion result to a second facial image adapted to the target age, to obtain a second fusion result; based on the second fusion result, obtaining the user's scenic spot photo result.
[0006] Furthermore, according to the age adjustment instruction, the first facial image displayed in the facial area in the first fusion result is adjusted to a second facial image adapted to the target age. The step of obtaining the second fusion result includes: extracting key facial feature points from the first facial image displayed in the facial area in the first fusion result; based on the key facial feature points, using a pre-trained age estimation model to predict the user's actual age; inputting the target age, actual age and first facial image carried in the age adjustment instruction into a pre-trained age change generation model, so as to output the second facial image adapted to the target age through the age change generation model, thereby obtaining the second fusion result.
[0007] Furthermore, based on the second fusion result, the step of obtaining the user's scenic spot photo result includes: updating the target theme template in the second fusion result so that the updated target theme template is adapted to the target age, and obtaining a third fusion result; and determining the third fusion result as the user's scenic spot photo result.
[0008] Furthermore, the age adjustment instruction sent by the user is generated in the following manner: multiple second facial images corresponding to the first facial image are displayed through a display interface; wherein each second facial image is adapted to a different age; and an image selection instruction of the second facial image adapted to the target age sent by the user is received to generate the age adjustment instruction according to the image selection instruction.
[0009] Furthermore, the age adjustment instruction sent by the user is generated by: displaying an age adjustment control through a display interface; and generating the age adjustment instruction in response to a first operation on the age adjustment control.
[0010] Furthermore, the method also includes: receiving a first production instruction sent by a user, and printing the results of photographing the scenic spot on a preset physical substrate according to the first production instruction to obtain a physical souvenir.
[0011] Furthermore, the method also includes: receiving a second production instruction sent by the user, collecting the user's facial three-dimensional data according to the second production instruction; and producing a 3D model souvenir corresponding to the user according to the facial three-dimensional data.
[0012] Furthermore, the method further includes: associating preset AR content with a target souvenir, so that the user can view the AR content through the target souvenir; wherein the target souvenir is a physical souvenir and / or a 3D model souvenir.
[0013] The present invention provides an artificial intelligence-based scenic spot photography device, which includes: a shooting module, which is used to receive a shooting instruction sent by a user and shoot a first facial image of the user according to the shooting instruction; a determination module, which is used to receive a theme template selection instruction sent by the user and determine a target theme template from multiple preset theme templates according to the theme template selection instruction; wherein the target theme template includes: a scenic spot theme, a virtual costume or a dynamic video template; a fusion module, which uses a preset artificial intelligence algorithm to fuse the first facial image into the facial area in the target theme template to obtain a first fusion result; and an acquisition module, which is used to obtain the user's scenic spot photography result based on the first fusion result; wherein the scenic spot photography result is: a scenic spot fused image or a scenic spot fused video.
[0014] The present invention provides an artificial intelligence-based scenic spot photography method and device. The method and device receive a photography instruction from a user and capture a first facial image of the user according to the photography instruction. The method also receives a theme template selection instruction from the user and, according to the theme template selection instruction, determines a target theme template from a plurality of preset theme templates. The target theme template includes at least one of the following: a scenic spot theme, a virtual costume, or a dynamic video template. The method utilizes a preset artificial intelligence algorithm to fuse the first facial image with the facial region in the target theme template to obtain a first fusion result. Based on the first fusion result, the method obtains a scenic spot photography result of the user. This method utilizes an artificial intelligence algorithm to fuse the first facial image of the user with the target theme template. Because the target theme template includes the scenic spot theme, virtual costume, dynamic video template, etc., the resulting scenic spot photography result can meet the user's personalized photography needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 A flowchart of a scenic spot photography method based on artificial intelligence provided by an embodiment of the present invention; Figure 2 A schematic structural diagram of a scenic spot photography device based on artificial intelligence provided by an embodiment of the present invention; Figure 3 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The technical solutions of the present invention are described clearly and completely below with reference to the embodiments. It is obvious that the embodiments described are only a portion of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without creative effort are also within the scope of protection of the present invention.
[0018] The existing technologies for photographing tourists at tourist attractions lack personalization, interactivity, convenience, affordability, depth of cultural integration, and impact on scenic area management and order. These methods struggle to meet tourists' growing demand for novel, personalized photography and their desire for in-depth cultural experiences. Specifically, existing technologies suffer from static image limitations, primarily focusing on static image generation and failing to meet user demands for dynamic videos and personalized experiences. Existing tourist photography methods lack interactivity, making it difficult to generate highly personalized images that deeply integrate with the culture or theme of a specific scenic area, and are unable to be customized to user needs. Traditional photography services and some commercial photography can disrupt scenic area tour order and the visitor experience. Tourists renting specialty clothing from scenic areas poses challenges, including high costs, cumbersome procedures, and clothing hygiene issues. Furthermore, the high costs of traditional photography services and rentals for specialty clothing limit the user's photography experience. Existing technologies struggle to deeply integrate users with the culture of the scenic area, failing to meet users' demand for in-depth cultural experiences. This lack of depth in cultural integration results in a poor cultural experience and sense of engagement for tourists at scenic areas. Based on this, an embodiment of the present invention provides a scenic spot photography method based on artificial intelligence, which can be applied to tourists taking photos in scenic spots.
[0019] To facilitate understanding of this embodiment, first, an artificial intelligence-based scenic spot photography method disclosed in an embodiment of the present invention is introduced. The method can be applied to an electronic device, which can specifically be an interactive travel photography device set in a scenic spot; Figure 1 As shown, the method includes the following steps: Step S102, receiving a photographing instruction sent by a user, and photographing a first facial image of the user according to the photographing instruction; The user can issue the above-mentioned photo-taking instruction by operating on the display screen of the electronic device, and the display screen can be a touch screen, etc. For example, a "photo-taking" control is displayed in the display interface corresponding to the display screen, and the user triggers the "photo-taking" control to issue the above-mentioned photo-taking instruction; after receiving the photo-taking instruction sent by the user, the user's first facial image can be captured by the built-in camera, and the first facial image of the user can be a frontal face image of the user, or a side face image, etc. The user can choose to capture the first facial image that meets the needs according to personal preference.
[0020] Step S104: receiving a theme template selection instruction sent by the user, and determining a target theme template from a plurality of preset theme templates according to the theme template selection instruction; wherein the target theme template includes at least one of the following: a scenic spot theme, a virtual costume, and a dynamic video template; In actual implementation, multiple theme templates can be pre-stored in the electronic device, and multiple theme templates can be displayed through the display interface. Each theme template can include its own corresponding scenic spot theme, virtual costumes, dynamic video templates, background templates, etc.; among them, the dynamic video template can be a real-shot video of the current scenic spot, a real-shot video of other scenic spots, a related video in a film and television work, a bird's-eye view video taken by a drone, etc.; for the related videos in the film and television works, in order to avoid infringement, the related videos in the film and television works can be stylized, for example, each frame of video image in the related video in the film and television work is converted into an ink painting style, a mosaic painting style, an oil painting style, etc.; the user can select the target theme template of interest from the multiple displayed theme templates, and after confirming the selected target theme template, the above-mentioned theme template selection instruction can be issued, and the electronic device can determine the target theme template selected by the user according to the theme template selection instruction.
[0021] Step S106, using a preset artificial intelligence algorithm to fuse the first facial image to the facial region in the target subject template to obtain a first fusion result; These AI algorithms can include deep learning, generative adversarial networks, and other AI algorithms. In practical implementation, the captured user's first facial image can be fused with the facial region of a target theme template to generate a new, personalized first fusion result that aligns with the theme. This first fusion result can be a static image or a dynamic video. This approach utilizes AI algorithms to achieve high-fidelity fusion of the two, ensuring a more natural and realistic integration of the person, the setting, and the costumes.
[0022] Step S108: Based on the first fusion result, obtain the user's scenic spot photo result.
[0023] In actual implementation, the first fusion result obtained above can be directly used as the user's scenic spot photo result. The user can also adjust the first fusion result according to actual needs. For example, the age corresponding to the first facial image displayed in the facial area of the first fusion result can be adjusted to adjust the first facial image, etc., and finally a scenic spot photo result that meets the user's needs is obtained. The scenic spot photo result can be a static scenic spot fusion image or a dynamic scenic spot fusion video.
[0024] The above-mentioned AI-based scenic spot photography method receives a photography instruction sent by a user and captures a first facial image of the user according to the photography instruction; receives a theme template selection instruction sent by the user and, according to the theme template selection instruction, determines a target theme template from a plurality of preset theme templates; wherein the target theme template includes at least one of the following: a scenic spot theme, a virtual costume, or a dynamic video template; utilizes a preset AI algorithm to fuse the first facial image with the facial region in the target theme template to obtain a first fusion result; and based on the first fusion result, obtains a scenic spot photography result of the user; wherein the scenic spot photography result is a scenic spot fusion image or a scenic spot fusion video. This method can utilize an AI algorithm to fuse the user's first facial image with the target theme template. Because the target theme template includes the scenic spot theme, virtual costume, dynamic video template, etc., the obtained scenic spot photography result can meet the user's personalized photography needs.
[0025] The embodiment of the present invention further provides another method for photographing a scenic spot based on artificial intelligence. The method is implemented on the basis of the method of the above embodiment, and the method includes the following steps: Step 1: receiving a photographing instruction sent by a user, and photographing a first facial image of the user according to the photographing instruction; Step 2: receiving a theme template selection instruction sent by the user, and determining a target theme template from a plurality of preset theme templates according to the theme template selection instruction; wherein the target theme template includes at least one of the following: a scenic spot theme, a virtual costume, and a dynamic video template; Step 3: Using a preset artificial intelligence algorithm, fuse the first facial image to the facial region in the target subject template to obtain a first fusion result; Step 4: Display the first fusion result through the display interface; In actual implementation, after the first fusion result is obtained, the first fusion result may be displayed in a display interface so that a user can preview the first fusion result.
[0026] Step 5: receiving an age adjustment instruction sent by the user, wherein the age adjustment instruction carries the target age; The above-mentioned target age may be the same as or different from the user's real age; in actual implementation, in the first fusion result, there may be a mismatch between the user's first facial image and the age of the corresponding virtual character in the target theme template. For example, the user's real age is 60 years old, and the age of the virtual character in the target theme template is 30 years old (such as "White Snake" in "The Legend of the New White Snake"). In this case, the user can operate the display interface to issue an age adjustment instruction carrying the target age.
[0027] In one implementation, the age adjustment instruction may be generated by following steps 50 and 51: Step 50: Displaying a plurality of second facial images corresponding to the first facial image via a display interface; wherein each second facial image is adapted to a different age; Step 51: receiving an image selection instruction for a second facial image adapted to a target age sent by a user, and generating an age adjustment instruction according to the image selection instruction.
[0028] In actual implementation, a first facial image can be used as a reference, and detailed facial features of the first facial image can be analyzed to determine a predicted age corresponding to the first facial image. This predicted age may be the same as or different from the user's actual age. For example, a user who takes good care of their face may be 50 years old, but the predicted age obtained by analyzing the captured first facial image is 40 years old. The detailed facial features of the first facial image can be used to predict the user's facial appearance, and multiple second facial images adapted for different ages can be obtained. These images can then be displayed on a display interface. For example, the predicted second facial images adapted for the user at ages 20, 30, and 60 can be displayed. The user can select a second facial image that meets their needs from the multiple displayed second facial images, such as a second facial image adapted for an age of 20, and so on, and then issue an image selection instruction. The electronic device can then generate a corresponding age adjustment instruction based on the image selection instruction.
[0029] In another implementation, the age adjustment instruction may also be generated by following steps 52 to 53: Step 52: displaying the age adjustment control via the display interface; Step 53: In response to the first operation on the age adjustment control, generate an age adjustment instruction.
[0030] The display form of the above-mentioned age adjustment control can be set according to actual needs. For example, it can be displayed in the form of a sliding bar, an edit box, etc.; for example, taking the age adjustment control in the form of a sliding bar as an example, the ages corresponding to the two ends of the sliding bar are 0 and 100 years old respectively. The user can drag the sliding control in the sliding bar. When the sliding control is located at different positions of the sliding bar, it corresponds to different ages. When the user slides the sliding control to the required position and confirms the position, the above-mentioned age adjustment instruction can be generated; for another example, taking the age adjustment control in the form of an edit box as an example, the user can directly enter the desired age in the edit box, or click the upward and downward triangle arrows corresponding to the edit box to adjust the age displayed in the edit box. After the user confirms the age in the edit box, the above-mentioned age adjustment instruction can be generated.
[0031] Step 6: According to the age adjustment instruction, the first facial image displayed in the facial region of the first fusion result is adjusted to a second facial image adapted to the target age, thereby obtaining a second fusion result; In actual implementation, after receiving the above-mentioned age adjustment instruction, the first facial image displayed in the facial area of the first fusion result can be adjusted according to the age adjustment instruction to adjust it to a second facial image adapted to the target age. The second facial image is fused with the target subject template in the first fusion result to obtain the above-mentioned second fusion result.
[0032] In one embodiment, the above step 6 may be implemented by following steps 60 to 62: Step 60, extracting key facial feature points from the first facial image displayed in the facial region in the first fusion result; The above-mentioned key facial feature points are usually mainly distributed in key positions of the facial features and facial contours such as the eyes, nose, and mouth; for example, the coordinates of the inner and outer corners of the eyes, the coordinates of the corners of the mouth, etc.; in actual implementation, face detection algorithms such as MTCNN (Multi-task Cascaded Convolutional Neural Networks) and RetinaFace (a face detection algorithm) can be used to locate the face in the first facial image and extract key facial feature points (such as eyes, nose, mouth, etc.).
[0033] Step 61 , predicting the user's actual age based on key facial feature points using a pre-trained age estimation model; The age estimation model can be a CNN (Convolutional Neural Network)-based model, etc. The age estimation model is typically trained based on a large number of age-labeled data sets. In actual implementation, after extracting key facial feature points from a first facial image, the first facial image can be input into the age estimation model to predict the user's actual age based on the key facial feature points in the first facial image.
[0034] Step 62: Input the target age, actual age, and first facial image carried in the age adjustment instruction into a pre-trained age change generation model, so as to output a second facial image adapted to the target age through the age change generation model to obtain a second fusion result.
[0035] In actual implementation, the above-mentioned target age, actual age and first facial image can be input into an age change generation model. The age change generation model can use a generative adversarial network (GAN) or other image generation technology to perform age change on the user's first facial image according to the target age to generate a second facial image adapted to the target age. For example, if the user wants to see what he or she would look like when he or she is younger or older, the age change generation model will adjust the facial features in the first facial image according to the target age to generate a corresponding second facial image, and fuse the second facial image with the target subject template in the first fusion result to ensure that the facial features in the second facial image are consistent with the character features in the target subject template to obtain the above-mentioned second fusion result.
[0036] Step seven: Based on the second fusion result, obtain the user's scenic spot photo result.
[0037] In one of the implementation methods, computer graphics technology can be used to render the second fusion result after fusion into a dynamic scenic spot fusion video. The specific steps include generating key frames, interpolation operations and the application of rendering engines, so that the playback of the scenic spot fusion video is smoother. In this embodiment, special effects such as filters and animation effects can also be added according to user instructions. The special effects library contains various filters, animation effects, etc. Users can choose appropriate special effects according to actual needs and apply them to the scenic spot fusion video. In one of the embodiments, as the scenic spot fusion video is played, the facial image of the user in the scenic spot fusion video can also show the effect of age change. For example, if the scenic spot fusion video shows a scene spanning many years, the age of the user in the scenic spot fusion video will also change accordingly. Correspondingly, the facial image of the user will also change with age. This embodiment can also optimize the age change algorithm corresponding to the age change generation model based on user feedback and data analysis results to enhance user experience.
[0038] For ease of understanding, the following provides an AI-powered scenic spot photography method. This method uses a scenic spot fusion video as an example. The specific implementation steps are as follows: 1. Data collection: Collect a large number of facial images of people of different ages and mark their age information.
[0039] 2. Model training: Use the collected dataset to train the age estimation model and the age change generation model. For example, convolutional neural networks such as ResNet and VGG can be used for training.
[0040] 3. Real-time processing: In actual applications, the user's first facial image is captured in real time and fused into the target subject template to obtain a first fusion result. Key facial feature points are extracted from the first facial image displayed in the facial area of the first fusion result, and the user's actual age is estimated using the trained age estimation model.
[0041] 4. Age change: Based on the target age specified by the user, the trained age change generation model is called to generate a second facial image with the corresponding age change.
[0042] 5. Fusion and rendering: The generated second facial image is fused with the target subject template and rendered into a dynamic video, i.e., the scenic spot fusion video.
[0043] 6. User interaction: Provide a user interface (corresponding to the above display interface) that allows users to select the target age, preview the effect and make adjustments.
[0044] Through the above steps, the age change algorithm can be implemented to provide users with a personalized photo-taking experience.
[0045] In one embodiment, step seven can be implemented by following steps 70 to 71: Step 70: Update the target subject template in the second fusion result so that the updated target subject template is adapted to the target age, thereby obtaining a third fusion result. Step 71: determine the third fusion result as the user's scenic spot photo result.
[0046] In actual implementation, when the electronic device confirms that the second facial image is compatible with the target theme template in the second fusion result, the second fusion result obtained above can be directly used as the user's scenic spot photo result; when it is confirmed that the second facial image is not compatible with the target theme template in the second fusion result, for example, the age corresponding to the second facial image is 20 years old, but the virtual clothing in the target theme template is the virtual clothing corresponding to the age of 40, then the target theme template in the second fusion result can be automatically updated, so that the updated target theme template can be more compatible with the target age, and finally a third fusion result is obtained, and the third fusion result can be used as the user's scenic spot photo result.
[0047] Step eight, receiving a first production instruction sent by the user, and printing the scenic spot photo result onto a preset physical substrate according to the first production instruction to obtain a physical souvenir.
[0048] The above-mentioned physical substrate can be a refrigerator magnet substrate, a postcard substrate, etc. In actual implementation, when a user needs to make a physical souvenir, the above-mentioned first production instruction can be issued. After the electronic device receives the first production instruction, it can use a thermal transfer process, UV (Ultraviolet LED Inkjet Printer, a full-color digital printing device that does not require platemaking) printing process, etc. to produce the scenic spot photo results on the physical substrate to obtain a prepared physical souvenir, such as a refrigerator magnet, postcard, etc. printed with a scenic spot fusion image. Among them, if the scenic spot photo result is a scenic spot fusion image, the scenic spot fusion image can be produced on the physical substrate. If the scenic spot photo result is a scenic spot fusion video, one or more frames of video images in the scenic spot fusion video can be selected, and the selected one or more frames of video images can be produced on the physical substrate.
[0049] Step nine, receiving a second production instruction sent by the user, and collecting the user's facial three-dimensional data according to the second production instruction; Step 10: Create a 3D model souvenir corresponding to the user based on the three-dimensional facial data.
[0050] The above-mentioned three-dimensional facial data may include information such as the shape, size, and position of parts such as the eyes, nose, mouth, and eyebrows; in actual implementation, users can also make 3D model souvenirs, such as personalized 3D figures, etc. At this time, the user can issue the above-mentioned second production instruction. After the electronic device receives the second production instruction, it can use a 3D printer or other molding technology to produce a personalized 3D figure, that is, a 3D model souvenir, according to the three-dimensional facial data; in another implementation method, 3D model souvenirs can also be made based on the above-mentioned scenic spot photography results.
[0051] Step 11: Associating the preset AR content with the target souvenir, so that the user can view the AR content through the target souvenir; wherein the target souvenir is a physical souvenir and / or a 3D model souvenir.
[0052] The above-mentioned AR content can be virtual performances of related stories, three-dimensional model displays, animation special effects, etc.; in actual implementation, AR technology can be used to empower the target souvenirs, that is, the target souvenirs produced are preset as AR markers, and users can subsequently use the corresponding mobile phone APP to scan the target souvenirs to watch or interactively experience the related AR content.
[0053] This AI-based scenic photography method enables dynamic video generation, age changes, clothing changes, background changes, multi-angle synthesis, and film and television plot imitation, meeting users' needs for a more vivid and personalized photography experience. Furthermore, by building a classification model and supporting engine, and implementing dynamic updates, it continuously improves the quality and efficiency of video generation, providing users with a more premium photography experience.
[0054] The following is another example of seamlessly embedding a user's first facial image into a background photo of a famous scenic spot, aiming to provide users with a richer and more personalized photo-taking experience. The specific solution can be as follows: 1. Background photo library establishment: 1) Collect and organize background photos of famous scenic spots, such as the Eiffel Tower, the Statue of Liberty, the Great Wall, etc.
[0055] 2) Classify and label background photos, including information such as location, style, and lighting conditions.
[0056] 3) Store the background photos on a server or local storage device and create an index for fast retrieval.
[0057] 2. User image collection: 1) Obtain the user image (corresponding to the first facial image mentioned above) through the camera or user upload.
[0058] 2) Preprocess user images, such as face detection and feature point extraction.
[0059] 3. Image fusion: 1) Based on user selection or system recommendation, select a suitable background photo (corresponding to the target theme template mentioned above) from the background photo library.
[0060] 2) Use image fusion algorithm to fuse the user image with the background photo.
[0061] Image fusion algorithms can be implemented based on the following technologies: a. Perspective Transformation: Based on the perspective relationship of the background photo, the user image is transformed to make it consistent with the perspective relationship of the background photo.
[0062] b. Color Correction: Adjust the color and brightness of the user's image to match the lighting conditions of the background photo.
[0063] c. Image stitching: Stitch the user image with the background photo to ensure that the images are connected naturally and there are no obvious stitching marks.
[0064] d. Generative Adversarial Network (GAN): Generate an image that is consistent with the style of the background photo using GAN and fuse it with the user’s image.
[0065] 4. Video generation (optional): 1) Using computer graphics technology, the fused images are rendered into dynamic videos.
[0066] 2) You can add animation effects, such as the user walking or turning around in the background.
[0067] 5. User Interaction 1) Provide a user interface that allows users to select a background photo, preview the effect, and make adjustments.
[0068] 2) Allow users to select different fusion algorithms and parameters to meet personalized needs.
[0069] 6. Result output: The final results are output to the display interface or storage device for users to view or share.
[0070] This embodiment seamlessly embeds user images into background photos of famous scenic spots, providing users with a richer and more personalized photo-taking experience. Users can easily place themselves in famous attractions around the world and capture memorable photos or videos.
[0071] In the above embodiment, a database containing background photos of well-known scenic spots can be established and managed to facilitate user selection and retrieval. Furthermore, a variety of image fusion techniques, such as perspective transformation, color correction, image stitching, and GAN, can be used to achieve seamless fusion of user images and background photos. Computer graphics technology can also be used to render the fused images into dynamic videos, enhancing the user experience. This embodiment also provides a user-friendly interactive interface that allows users to select background photos, preview the effects, and make adjustments to meet personalized needs.
[0072] For ease of understanding, the following describes an interactive travel photography device. First, the hardware components of the device are described, which specifically include the following hardware: 1. Main structure: Typically a vertical, integrated design, consisting of a metal frame (e.g., 1.5mm galvanized steel) and an outer shell (e.g., powder-coated). The printer typically integrates a printing unit and is relatively heavy (e.g., approximately 280 kg).
[0073] 2. Interactive Display Module: A vertically oriented touchscreen display for user interface display, information input, and interactive operations. It features a 55-inch TFT-LCD (Thin Film Transistor Liquid Crystal Display) with a 4K resolution (3840x2160), high brightness (e.g., 2500cd / m2), 10-point G+G capacitive touch, and a 6mm AR anti-glare tempered glass surface.
[0074] 3. Image acquisition module: built-in high-definition camera, used to capture the user's front or half-body facial image. 4. Processing module: This module has a built-in high-performance computing unit (e.g., an Intel Core i5-10400 processor, 16GB of RAM, 128GB of storage, and runs Ubuntu). Alternatively, it may be connected to a cloud server via a network and is responsible for running the user interface, processing images, executing AI algorithms, and controlling peripherals.
[0075] 5. Output module (optional): Photo Printer: This device is used to print high-quality AI-generated photos (corresponding to the aforementioned scenic spot images), allowing users to print their own photos. In practical applications, multiple (e.g., four) photo printers can be integrated. Multiple printers improve printing efficiency, increase the concurrent processing capability and reliability of photo output, achieve load balancing, and provide redundant backup. Photo printers typically have a photo exit for users to remove printed photos.
[0076] Built-in or external production unit: used to produce physical souvenirs such as refrigerator magnets and postcards. For example, it may include: refrigerator magnet production equipment: for example, using thermal transfer, UV printing or other methods to produce AI-generated images on refrigerator magnet substrates.
[0077] 3D figure production equipment: For example, a 3D printer or other molding technology can be used to generate photos or three-dimensional model data combined with the user's image based on AI (corresponding to the above-mentioned three-dimensional facial data) to produce personalized 3D figures (corresponding to the above-mentioned 3D model souvenirs).
[0078] 6. Auxiliary modules: such as ambient light sensor, speaker (8Ω 10W2), network module (LAN / WiFi / 5G), cooling system (such as intelligent temperature control air cooling), necessary interfaces (such as USB2.02, USB3.0*1), enhanced cooling system (such as 1500W refrigeration industrial air conditioner and fan), etc.
[0079] 7. Lighting system: such as acrylic light boxes, mini luminous characters, embedded light strips, etc., which can adjust the lighting effects, create an atmosphere and attract users.
[0080] 8. Audio system: used to play prompt sounds or background music.
[0081] 9. Network module: used to connect to cloud servers for AI calculations, data updates, remote management, etc. The above-mentioned interactive travel photography equipment has good environmental adaptability. Specifically, it can have an IP65 protection level, waterproof, dustproof, anti-theft, explosion-proof, and rust-proof design, and support a wider operating temperature range (such as -20℃-50℃, and can also have a heater function). It adopts an industrial-grade air-conditioning cooling system and a wide temperature design to ensure that the device can operate stably for a long time in unattended outdoor or semi-outdoor environments.
[0082] The following is an explanation of the device's software system and core algorithms: 1. User Interface (UI): Based on the touch screen, it guides users to complete operations such as taking photos, selecting themes, previewing effects, making payments, and selecting souvenirs.
[0083] 2. Face detection and feature extraction: After the camera captures the image, the system identifies the face area and extracts key facial feature points.
[0084] 3. AI image generation engine: Receives the user's facial feature data (corresponding to the first facial image mentioned above) and the selected target theme template (such as clothing, makeup, background, dynamic video template, composition integrated with the scenery, etc. for a specific scenic spot).
[0085] Leveraging AI algorithms like deep learning and generative adversarial networks (GANs), users' facial data is integrated with a target theme template for high-fidelity fusion, generating new, personalized photos or videos that align with the theme (corresponding to the aforementioned scenic spot photography results). These optimized AI algorithms ensure a more natural and realistic integration of the subject and the setting / costume.
[0086] 4. AR enhancement module (optional): Associate specific AR content with generated physical souvenirs (such as postcards).
[0087] When users scan physical souvenirs through a mobile phone APP, they can trigger preset AR effects, such as virtual interpretation of related stories, three-dimensional model displays, animation special effects, etc., to enhance the interactivity and fun of the physical souvenirs.
[0088] 5. Cloud service platform (optional): Store and manage digital cultural asset libraries (such as theme templates and clothing models of various scenic spots).
[0089] Provides cloud-based AI computing services to share or bear the computing pressure of local devices.
[0090] Support remote system updates and maintenance.
[0091] Conduct data statistics and analysis.
[0092] In particular: the device must support the transmission of the generated image data or the corresponding 3D model data to the control system of the external production unit; The device must include a complete self-service process, including payment, printing instruction sending, printing status monitoring, and consumables remaining monitoring (optional).
[0093] It has print task management function and can distribute print requests to multiple internal printers.
[0094] The following describes the working process of the equipment: a. The user initiates the experience via a touch screen.
[0095] b. The device guides the user to follow the instructions and uses the built-in camera to capture the user's first facial image.
[0096] c. The user selects a target theme template of interest on the screen. The target theme template may include: scenic theme, virtual clothing, dynamic video template or background template, etc.
[0097] d. Send the user’s first facial image and the selected target subject template to the AI image generation engine.
[0098] e. The AI image generation engine processes the first facial image and the selected target subject template to generate a fused personalized photo or video, which is previewed on the screen.
[0099] f. After confirming the result, the user can choose to pay and receive a digital photo, or choose to create a physical souvenir (photo, postcard, refrigerator magnet, etc.).
[0100] g. If you select a physical souvenir, the device's output module (photo printer / production unit) will produce the corresponding item according to your instructions. Production time varies depending on the item type (e.g., approximately 10-30 seconds for a photo, 3-5 minutes for a refrigerator magnet).
[0101] h. If the device includes 3D printing capabilities, it can extract the user's 3D facial data and quickly create 3D models based on that data. This allows for the 3D printing of physical souvenirs such as a 3D model (fig) of the user.
[0102] i. If the product includes AR functionality, the physical souvenir will be pre-set with an AR marker. Users can then scan the souvenir using the accompanying mobile app to view or interact with the associated AR content. The AI-based scenic spot photography method described above aims to provide an interactive travel photography method that integrates artificial intelligence (AI), augmented reality (AR), and 3D printing technologies, as well as corresponding methods for generating and enhancing photos and videos. This method can capture a user's facial image and leverage AI technology to quickly generate personalized photos or videos that are highly consistent with the scenic spot's characteristics or theme (e.g., virtual costume changes, scene integration). The photos can also be transformed into physical souvenirs (e.g., postcards, refrigerator magnets, etc.). Furthermore, 3D printing and AR technologies can be used to empower physical souvenirs, providing a virtual interactive experience, thereby enhancing the user experience, creating unique cultural commemorative value, and providing organizers with an innovative service and management approach.
[0103] The above-mentioned scenic spot photography method based on artificial intelligence can produce the following beneficial effects: 1. AI-driven scenario-based portrait fusion system: The core lies in AI image generation technology that combines specific tourist scenes (scenic area culture, historical costumes, and characteristic elements). It can quickly and with high fidelity integrate the user's face into the preset theme, achieving personalized photo or video effects such as "virtual dress-up" or "scene travel".
[0104] 2. Diversified souvenir generation system: Combining the AI interactive kiosk (self-service) with a multi-functional production unit, it can not only generate AI photos, but also produce a variety of complex or specialized personalized physical souvenirs such as refrigerator magnets and 3D figurines. 3. When the production unit is external, this architecture, which separates the main interaction kiosk from the external production unit (or loosely couples it), facilitates flexible configuration of production capacity based on site and business needs. The kiosk focuses on user interaction and AI generation.
[0105] 4. Travel photography service processes for high-value-added customized products enable the customized production of personalized physical products (such as refrigerator magnets and 3D figurines). Human participation in the production process allows for customized production and quality control of personalized products with higher craftsmanship requirements and greater complexity. AI-generated personalized physical souvenirs (such as postcards and 3D-printed figurines) can be used as AR markers, and scanning can trigger virtual content (such as storytelling and 3D models). This combines physical mementos with digital experiences, enhancing the value and interest of souvenirs. Emerging technologies (such as AI) can be used to enhance visitors' cultural experience and sense of participation in scenic spots, creating new value.
[0106] 5. Travel photography service method based on cloud-based digital cultural asset library: Utilize cloud platforms to store and manage rich digital cultural materials (templates, models, etc.) related to local scenic spots, and provide content and AI computing support to travel photography devices deployed in various locations through the network, achieving rapid content updates and scalability of services.
[0107] 6. Travel photography solutions for tourist experience and scenic spot management: This approach not only provides tourists with novel interactive experiences and personalized souvenirs, but also provides scenic spots with a digital solution to manage tourist photography behavior, reduce disruptions, spread culture, and increase revenue.
[0108] 7. This method can integrate AI interaction, image generation, and multiple photo printers into a single self-service device, enabling users to complete the process of personalized photo generation and printing by themselves. 8. This method simplifies the self-service photo-taking service process and provides a clear user interface and guidance. Users can independently complete the entire process from taking photos, selecting themes, generating previews to payment and pickup without manual intervention.
[0109] 9. This method uses artificial intelligence algorithms to fuse the user's facial image into a target theme template, enabling a richer, more personalized photography experience, such as video generation, age change, clothing change, background change, multi-angle synthesis, and film and television plot imitation. This method breaks through the limitations of traditional static images and enables the generation of dynamic videos, providing users with a more vivid and immersive experience.
[0110] 10. Support users to choose different theme templates according to their own preferences, and carry out multi-dimensional customization such as age, clothing, background, multiple angles, film and television plots, etc. to meet the diverse needs of users.
[0111] 11. Use self-service photography equipment and AI technology to reduce photography costs and provide multiple payment methods, such as QR code payment and online payment.
[0112] 12. Deep Application of AI Technology: Leveraging AI technologies such as deep learning and generative adversarial networks, we achieve high-fidelity image fusion and video rendering. We can deeply integrate user facial images with the cultural context of scenic spots. For example, we can integrate user facial images into scenic spot theme templates or generate videos that mimic film and television plots, enhancing users' cultural experience and sense of engagement, providing users with a more vivid, personalized, and immersive photography experience.
[0113] 13. Dynamic update: Dynamic update of classification models and supporting engines can be achieved through cloud servers to ensure the advancement and adaptability of technology.
[0114] 14. Multi-language support: It can support multiple language interfaces, which is convenient for users of different languages.
[0115] 15. User feedback mechanism: It can collect user feedback and optimize algorithms to continuously improve user experience.
[0116] 16. Security and privacy protection: Technical means such as data encryption and permission control can be used to ensure the security and privacy of user data.
[0117] 17. Business model innovation: A variety of business models can be provided, such as payment models, advertising placement, cooperative promotion, etc., to realize commercial value.
[0118] 18. Wide range of application scenarios: It can be applied to multiple fields such as scenic spot photography, tourism platforms, social media, film and television shooting, virtual idols, education and entertainment, etc., bringing users a new photography experience and promoting the development of related industries; it can create new value for partners such as scenic spots, tourism platforms, and film and television companies; it can promote cultural communication and education by simulating historical scenes, scientific experiments, and famous attractions around the world.
[0119] In short, this solution has broad application prospects and market potential, will bring a new photography experience to users, and promote the development of related industries.
[0120] The embodiment of the present invention provides a scenic spot photographing device based on artificial intelligence, such as Figure 2 As shown, the device includes: a shooting module 20, which is used to receive a shooting instruction sent by a user and shoot a first facial image of the user according to the shooting instruction; a determination module 21, which is used to receive a theme template selection instruction sent by the user and determine a target theme template from a plurality of preset theme templates according to the theme template selection instruction; wherein the target theme template includes at least one of the following: a scenic spot theme, a virtual costume, and a dynamic video template; a fusion module 22, which is used to use a preset artificial intelligence algorithm to fuse the first facial image into the facial area in the target theme template to obtain a first fusion result; an acquisition module 23, which is used to obtain a scenic spot shooting result of the user based on the first fusion result; wherein the scenic spot shooting result is: a scenic spot fused image or a scenic spot fused video.
[0121] The above-mentioned artificial intelligence-based scenic spot photography device can use artificial intelligence algorithms to fuse the user's first facial image into the target theme template. Since the target theme template contains the scenic spot theme, virtual clothing, dynamic video template, etc., the obtained scenic spot photography results can meet the user's personalized photography needs.
[0122] Furthermore, the acquisition module is also used to: display the first fusion result through the display interface; receive an age adjustment instruction sent by the user, wherein the age adjustment instruction carries the target age; according to the age adjustment instruction, adjust the first facial image displayed in the facial area of the first fusion result to a second facial image adapted to the target age to obtain a second fusion result; based on the second fusion result, obtain the user's scenic spot photo result.
[0123] Furthermore, the acquisition module is also used to: extract key facial feature points from the first facial image displayed in the facial area in the first fusion result; based on the key facial feature points, use a pre-trained age estimation model to predict the user's actual age; input the target age, actual age and first facial image carried in the age adjustment instruction into a pre-trained age change generation model, so as to output a second facial image adapted to the target age through the age change generation model to obtain a second fusion result.
[0124] Furthermore, the acquisition module is also used to: update the target theme template in the second fusion result so that the updated target theme template is adapted to the target age, thereby obtaining a third fusion result; and determine the third fusion result as the user's scenic spot photo result.
[0125] Furthermore, the age adjustment instruction sent by the user is generated in the following manner: multiple second facial images corresponding to the first facial image are displayed through a display interface; wherein each second facial image is adapted to a different age; and an image selection instruction of the second facial image adapted to the target age sent by the user is received to generate the age adjustment instruction according to the image selection instruction.
[0126] Furthermore, the age adjustment instruction sent by the user is generated by: displaying an age adjustment control through a display interface; and generating the age adjustment instruction in response to a first operation on the age adjustment control.
[0127] Furthermore, the device is also used to: receive a first production instruction sent by a user, and print the results of photographing the scenic area onto a preset physical substrate according to the first production instruction to obtain a physical souvenir.
[0128] Furthermore, the device is also used to: receive a second production instruction sent by the user, collect the user's facial three-dimensional data according to the second production instruction; and produce a 3D model souvenir corresponding to the user according to the facial three-dimensional data.
[0129] Furthermore, the device is also used to: associate preset AR content with a target souvenir, so that the user can view the AR content through the target souvenir; wherein the target souvenir is a physical souvenir and / or a 3D model souvenir.
[0130] The artificial intelligence-based scenic spot photographing device provided in the embodiment of the present invention has the same implementation principle and technical effects as the aforementioned artificial intelligence-based scenic spot photographing method embodiment. For the sake of brief description, for any parts not mentioned in the embodiment of the artificial intelligence-based scenic spot photographing device, please refer to the corresponding content in the aforementioned artificial intelligence-based scenic spot photographing method embodiment.
[0131] The embodiment of the present invention further provides an electronic device, see Figure 3As shown, the electronic device includes a processor 130 and a memory 131. The memory 131 stores machine-executable instructions that can be executed by the processor 130. The processor 130 executes the machine-executable instructions to implement the above-mentioned artificial intelligence-based scenic spot photography method.
[0132] Furthermore, Figure 3 The electronic device shown further includes a bus 132 and a communication interface 133 , and the processor 130 , the communication interface 133 and the memory 131 are connected via the bus 132 .
[0133] The memory 131 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage. The communication connection between the system network element and at least one other network element is achieved through at least one communication interface 133 (which may be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. may be used. The bus 132 may be an ISA bus, a PCI bus, or an EISA bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0134] The processor 130 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 130 or by software instructions. The above processor 130 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present invention can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 131, and processor 130 reads information in memory 131 and, in conjunction with its hardware, completes the steps of the method of the aforementioned embodiment.
[0135] An embodiment of the present invention also provides a machine-readable storage medium, which stores machine-executable instructions. When the machine-executable instructions are called and executed by the processor, the machine-executable instructions prompt the processor to implement the above-mentioned artificial intelligence-based scenic spot photography method. The specific implementation can be found in the method embodiment, which will not be repeated here.
[0136] The computer program product of the artificial intelligence-based scenic spot photography method and device provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the previous method embodiments. For specific implementation, please refer to the method embodiments and will not be repeated here.
[0137] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A scenic spot photography method based on artificial intelligence, characterized in that: The method comprises: receiving a photographing instruction sent by a user, and photographing a first facial image of the user according to the photographing instruction; receiving a theme template selection instruction sent by the user, and determining a target theme template from a plurality of preset theme templates according to the theme template selection instruction; wherein the target theme template includes at least one of the following: a scenic spot theme, a virtual costume, and a dynamic video template; Using a preset artificial intelligence algorithm, fuse the first facial image to the facial region in the target subject template to obtain a first fusion result; Based on the first fusion result, a scenic spot photo result of the user is obtained; wherein the scenic spot photo result is: a scenic spot fusion image or a scenic spot fusion video.
2. The method according to claim 1, characterized in that The step of obtaining a user's scenic spot photo result based on the first fusion result includes: Displaying the first fusion result through a display interface; receiving an age adjustment instruction sent by the user, wherein the age adjustment instruction carries a target age; adjusting, according to the age adjustment instruction, the first facial image displayed in the facial region in the first fusion result to a second facial image adapted to the target age, to obtain a second fusion result; Based on the second fusion result, a scenic spot photo result of the user is obtained.
3. The method according to claim 2, characterized in that The step of adjusting, according to the age adjustment instruction, the first facial image displayed in the facial region in the first fusion result to a second facial image adapted to the target age, to obtain the second fusion result comprises: extracting key facial feature points from the first facial image displayed in the facial region in the first fusion result; Predicting the actual age of the user using a pre-trained age estimation model based on the key facial feature points; The target age, the actual age, and the first facial image carried in the age adjustment instruction are input into a pre-trained age change generation model, so that a second facial image adapted to the target age is output through the age change generation model to obtain a second fusion result.
4. The method according to claim 2, characterized in that The step of obtaining the user's scenic spot photo result based on the second fusion result includes: Updating the target subject template in the second fusion result so that the updated target subject template is adapted to the target age, thereby obtaining a third fusion result; The third fusion result is determined as the scenic spot photo-taking result of the user.
5. The method according to claim 2, characterized in that The age adjustment instruction sent by the user is generated in the following manner: Displaying a plurality of second facial images corresponding to the first facial image through the display interface; wherein each of the second facial images is adapted to a different age; An image selection instruction for a second facial image adapted for a target age is received from the user, so as to generate the age adjustment instruction according to the image selection instruction.
6. The method according to claim 2, characterized in that The age adjustment instruction sent by the user is generated in the following manner: displaying an age adjustment control via the display interface; The age adjustment instruction is generated in response to a first operation on the age adjustment control.
7. The method according to claim 1, characterized in that The method further comprises: A first production instruction sent by the user is received, and according to the first production instruction, the photographic result of the scenic spot is printed on a preset physical substrate to obtain a physical souvenir.
8. The method according to claim 7, characterized in that The method further comprises: receiving a second production instruction sent by the user, and collecting three-dimensional facial data of the user according to the second production instruction; A 3D model souvenir corresponding to the user is produced according to the three-dimensional facial data.
9. The method according to claim 8, characterized in that The method further comprises: The preset AR content is associated with a target souvenir, so that the user can view the AR content through the target souvenir; wherein the target souvenir is the physical souvenir and / or the 3D model souvenir.
10. A scenic spot photography device based on artificial intelligence, characterized in that: The device comprises: A shooting module is used to receive a shooting instruction sent by a user and shoot a first facial image of the user according to the shooting instruction; a determination module, configured to receive a theme template selection instruction sent by the user, and determine a target theme template from a plurality of preset theme templates according to the theme template selection instruction; wherein the target theme template includes at least one of the following: a scenic spot theme, a virtual costume, and a dynamic video template; a fusion module, configured to fuse the first facial image with the facial region in the target subject template using a preset artificial intelligence algorithm to obtain a first fusion result; An acquisition module is used to obtain the user's scenic spot photo result based on the first fusion result; wherein the scenic spot photo result is: a scenic spot fusion image or a scenic spot fusion video.