Scenic spot tourist video snapshot processing method
By adopting facial recognition technology and cloud deployment methods in scenic spots, automatic photo shooting and video editing is achieved, which solves the problem that tourists find it difficult to capture ideal photos and videos, and improves user satisfaction and stickiness.
Patent Information
- Application Number
- CN202510213234.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
It is difficult for tourists to capture ideal photos and videos during the trip, and the post-editing process is time-consuming and labor-intensive, complicating the process of recording travel memories.
A method of video capture processing for scenic spot tourists is adopted, and face recognition is performed through the InsightFace library combining CUDA and TensorRT, and face detection and recognition is performed by combining RetinaFace, ArcFace and EfficientNet to realize automatic photo shooting and video editing. The method includes steps such as image acquisition and training, face detection and recognition, cloud deployment, editing and synthesis, user interaction and user sharing.
It realizes automatic photo shooting and video editing, provides personalized and customized services, increases user satisfaction and stickiness, and simplifies the process of tourists recording travel memories.
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of scenic spot photography, and in particular to a method for capturing and processing videos of tourists in scenic spots. Background Art
[0002] As an important part of the global economy, the tourism industry has experienced rapid development in recent years. With the advancement of information technology, especially the popularity of the Internet and mobile devices, the way the tourism industry operates has undergone profound changes. However, tourists often face some headaches during their travels, such as difficulty in capturing and editing. In a busy travel environment, tourists hope to capture wonderful moments, but often due to time constraints or environmental factors, it is difficult to take ideal photos. In addition, the process of post-editing videos or photos is often time-consuming and labor-intensive, making it more complicated to record travel memories. Summary of the invention
[0003] The purpose of the present invention is to solve the shortcomings of the prior art and propose a method for capturing and processing tourist videos in a scenic spot. The technical solution adopted by the present invention is:
[0004] A method for capturing and processing video of tourists in a scenic spot, the method comprising the following steps:
[0005] S1, technology selection and component selection;
[0006] S2, image acquisition and image training;
[0007] S3, face detection and face recognition;
[0008] S4, cloud deployment;
[0009] S5, editing and synthesis;
[0010] S6. User interaction;
[0011] S7. User sharing.
[0012] As an improvement, the specific steps of S1, technology selection and component selection are as follows: In terms of face recognition technology, the InsightFace library is selected, combined with CUDA and TensorRT for inference. Video slices are sent to algorithm parsing through rtmp real-time streaming. The algorithm makes a preliminary slice anomaly judgment, and then sends it to deepstream for video stream parsing. Deepstream combines the retinaface face detection model provided by InsightFace to return the detected face area, and then uses Arcface and EfficientNet for post-processing. The combination of the two gives full play to the performance of the GPU, ensuring the efficient operation of the model in the training and inference stages, and supporting the real-time processing and analysis of large-scale video data.
[0013] As an improvement, the specific steps of S2, image acquisition and image training are as follows: In the image acquisition stage, to ensure the diversity and quality of data, we use high-resolution cameras, combined with the feature extraction ability of the algorithm, to capture facial images from different angles and lighting conditions in real time, and preprocess the captured images, including face processing and data augmentation. Using the technology mentioned in S1, image training is carried out. After the training is completed, the final model is obtained by evaluating and optimizing various indicators such as the accuracy and precision of the model.
[0014] As an improvement, the specific steps of S3, face detection and face recognition are as follows: This system uses RetinaFace as the face detection module, ArcFace as the face recognition module, and EfficientNet for facial expression recognition.
[0015] As an improvement, the specific steps of S4, cloud deployment are as follows: Select a suitable cloud platform, configure computing resources and storage services according to needs, use containerization technology (Docker) to deploy the application environment, upload the trained model to the cloud, configure the corresponding API interface, so that it can receive video data and return results in real time. We divide the algorithm into two parts: the front end is responsible for face detection and feature extraction, and the back end focuses on face recognition and face comparison. The advantage of this design is that it can effectively share the computing load, enabling the front end and the back end to perform their respective duties. The front end can quickly identify and process user requests, while the back end can use stronger computing resources for in-depth analysis, reducing the burden on the back-end algorithm service and ensuring the stable operation of the system under high concurrency.
[0016] As an improvement, the specific steps of S5, video clip synthesis are as follows: Based on the face information recognized in S4, determine the video clips to be clipped, record the start and end times of each video clip, and send the video clips to the clip server through the API for video clip clipping. On the clip server, clip the video clips according to pre-set rules (such as clip style, MP4, etc.), and you can choose to add elements such as transition effects and subtitles. Synthesize the clipped clips into a complete Vlog video, and store the generated video file in the cloud storage service for subsequent access.
[0017] As an improvement, the specific steps of S6, user interaction are as follows: Based on the functions implemented in S5, generate an accessible link (usually a URL) for the stored video file. The user side designs a video playback interface. The user first scans the scenic area channel code, and through the face recognition interface designed by the small program, obtains the wonderful clips during the play. Our big data center analyzes the tourist's play footprint and play preferences, invokes the AI Doubao large model, interacts with the AI to obtain the user's needs, and recommends play strategies for the user.
[0018] As an improvement, the specific steps of S7, user sharing are as follows: The user obtains their own play video and play strategy. Our small program provides a sharing function, which can share their short videos to major short video platforms, drain traffic from each platform, perform video mounting, and obtain corresponding commissions through consumption. The next section of the sharing section is to call the AI dressing and AI drawing developed by us to support users to customize videos, perform secondary clipping, and drain traffic after publishing.
[0019] The beneficial effects of the present invention are as follows:
[0020] A method for processing video capture of scenic area tourists in the present invention automatically takes pictures through the cameras in the scenic area in an artificial intelligence manner, forms a personal account for face recognition, and then processes the pictures captured by the cameras through AI to form the required photos. In this way, the present invention automatically provides personalized and customized services for individuals, enabling consumers to better obtain a sense of satisfaction, thereby increasing user stickiness. Specific implementation mode
[0021] In order to make the content of the present invention easier to be clearly understood, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention.
[0022] The method for processing video capture of scenic area tourists includes the following steps:
[0023] S1. Technology and Component Selection: In terms of face recognition technology, the InsightFace library was selected, combined with CUDA and TensorRT for inference. Video slices were sent to the algorithm for parsing through real-time RTMP streaming. The algorithm performed preliminary slice anomaly detection and then sent the data to DeepStream for video stream parsing. DeepStream combined with the retinaface face detection model provided by InsightFace to return the detected face regions. Then, Arcface and EfficientNet were used for post-processing. The combination of the two fully utilized the performance of the GPU, ensuring the efficient operation of the model during the training and inference phases and supporting the real-time processing and analysis of large-scale video data.
[0024] S2. Image Acquisition and Image Training: In the image acquisition phase, to ensure the diversity and quality of the data, we used high-resolution cameras. Combining with the feature extraction ability of the algorithm, facial images from different angles and lighting conditions were captured in real-time. The captured images were preprocessed, including face processing and data augmentation. Using the technology mentioned in S1, image training was carried out. After the training was completed, the final model was obtained by evaluating and optimizing various indicators such as the accuracy and precision of the model.
[0025] S3. Face Detection and Face Recognition: This system uses RetinaFace as the face detection module, ArcFace as the face recognition module, and EfficientNet for facial expression recognition.
[0026] S4. Cloud Deployment: Select a suitable cloud platform, configure computing resources and storage services according to needs, deploy the application environment using containerization technology (Docker), upload the trained model to the cloud, configure the corresponding API interface, so that it can receive video data and return results in real-time. We divide the algorithm into two parts: the front-end is responsible for face detection and feature extraction, and the back-end focuses on face recognition and face comparison. The advantage of this design is that it can effectively share the computing load, enabling the front-end and back-end to perform their respective functions. The front-end can quickly identify and process user requests, while the back-end can use stronger computing resources for in-depth analysis, reducing the burden on the back-end algorithm service and ensuring the stable operation of the system under high concurrency.
[0027] S5. Editing and synthesis: Based on the face information recognized in S4, determine the video segments to be edited, record the start and end times of each video segment, and send the video segments to the editing server via the API for video segment editing. On the editing server, edit the video segments according to preset rules (such as editing style, MP4, etc.), and you can choose to add elements such as transition effects and subtitles. Synthesize the edited segments into a complete Vlog video and store the generated video file in the cloud storage service for subsequent access;
[0028] S6. User interaction: Based on the functions implemented in S5, generate an accessible link (usually a URL) for the stored video file. The user side is designed with a video playback interface. The user first scans the scenic area channel code and obtains the wonderful segments during the play through the face recognition interface designed by the small program. Our big data center analyzes the tourist's play footprint and play preferences, invokes the AI Doubao large model, interacts with the AI to obtain the user's needs, and recommends play strategies for the user;
[0029] S7. User sharing: After the user obtains their own play video and play strategy, our small program provides a sharing function, which can share their short video to major short video platforms, attract traffic from each platform, perform video mounting, and obtain corresponding commissions through consumption. The next section of the sharing section is to call the AI clothing change and AI drawing we developed to support users to customize videos, perform secondary editing, and attract traffic after publishing.
[0030] In an artificial intelligence way, automatically take pictures through the cameras in the scenic area, form a personal account for face recognition, and then process the pictures captured by the cameras through AI to form the required photos. In this way, the present invention automatically provides personalized and customized services for individuals, enabling consumers to better obtain a sense of satisfaction, thereby increasing user stickiness.
[0031] The above are only the preferred embodiments of this invention patent and are not intended to limit this invention patent. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this invention patent shall be included within the protection scope of this invention patent.
Claims
1. A method for capturing and processing video of tourists in a scenic spot, characterized in that: The method for capturing and processing video of tourists in scenic spots includes the following steps: S1, technology selection and component selection; S2, image acquisition and image training; S3, face detection and face recognition; S4, cloud deployment; S5, editing and synthesis; S6. User interaction; S7. User sharing.
2. A method for capturing and processing tourist video in a scenic spot according to claim 1, characterized in that: The specific steps of S1, technology selection and component selection are as follows: in terms of face recognition technology, the InsightFace library is selected, combined with CUDA and TensorRT for reasoning, and video slices are sent to the algorithm for analysis through rtmp real-time streaming. The algorithm makes a preliminary slice anomaly judgment and then sends deepstream for video stream analysis. Deepstream combines the retinaface face detection model provided by InsightFace to return the face area of the detection result, and then uses Arcface and EfficientNet for post-processing. The combination of the two gives full play to the performance of the GPU, ensuring the efficient operation of the model in the training and reasoning stages, and supports real-time processing and analysis of large-scale video data.
3. A method for capturing and processing tourist video in a scenic spot according to claim 1, characterized in that: The specific steps of S2, image acquisition and image training are as follows: in the image acquisition stage, in order to ensure the diversity and quality of the data, we used a high-resolution camera, combined with the feature extraction capability of the algorithm, to capture facial images from different angles and lighting in real time, and pre-processed the collected images, including face processing and data enhancement, and used the technology mentioned in S1 to perform image training. After the training is completed, the model's accuracy, precision and other indicators are evaluated and optimized to obtain the final model.
4. A method for capturing and processing tourist video in a scenic spot according to claim 1, characterized in that: The specific steps of S3, face detection and face recognition, are as follows: this system uses RetinaFace as the face detection module, ArcFace as the face recognition module, and EfficientNet for facial expression recognition.
5. A method for capturing and processing tourist video in a scenic spot according to claim 1, characterized in that: The specific steps of S4, cloud deployment are: select a suitable cloud platform, configure computing resources and storage services as needed, use containerization technology (Docker) to deploy the application environment, upload the trained model to the cloud, configure the corresponding API interface, so that it can receive video data and return results in real time. We divide the algorithm into two parts: the front end is responsible for face detection and feature extraction, and the back end focuses on face recognition and face comparison. The advantage of this design is that it can effectively share the computing load, so that the front end and the back end can perform their respective duties. The front end can quickly identify and process user requests, while the back end can use stronger computing resources for in-depth analysis, reducing the burden on the back end algorithm service and ensuring that the system can still run stably under high concurrency.
6. A method for capturing and processing tourist video in a scenic spot according to claim 1, characterized in that: The specific steps of S5, editing and synthesis are: based on the facial information recognized by S4, determine the video clips that need to be edited, record the start and end time of each video clip, send the video clips to the editing server through the API for editing, and on the editing server, edit the video clips according to pre-set rules (such as editing style, MP4, etc.), and you can choose to add transition effects, subtitles and other elements, synthesize the edited clips into a complete Vlog video, and store the generated video files in the cloud storage service for subsequent access.
7. A method for capturing and processing tourist video in a scenic spot according to claim 1, characterized in that: The specific steps of S6 and user interaction are: based on the function implemented in S5, an accessible link (usually a URL) is generated for the stored video file. The user end designs a video playback interface. The user first scans the scenic spot channel code and obtains the wonderful clips of the tour through the face recognition interface designed by the mini program. Our big data center analyzes the tourists' travel footprints and travel preferences, retrieves the AI Bean Bag model, interacts with AI to obtain user needs, and recommends travel strategies for users.
8. A method for capturing and processing tourist video in a scenic spot according to claim 1, characterized in that: The specific steps of S7 and user sharing are as follows: users obtain their own game videos and game strategies. Our mini program provides a sharing function, which allows users to share their short videos to major short video platforms, attract traffic from various platforms, mount videos, and get corresponding commissions through consumption. The next section of the sharing section is to call our developed AI dressing, AI drawing, etc. to support user-defined videos, perform secondary editing, and attract traffic after release.