Scene material library-based photographing composition generation and photographing prompting system
Through an intelligent photography composition generation and photography prompt system based on the scene material library, the problems of poor adaptability of posture and scene, low interaction efficiency and insufficient generation stability in the existing technology are solved, efficient and accurate shooting guidance and dynamic shooting are achieved, and user experience and shooting quality are improved.
Patent Information
- Application Number
- CN202510458877.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
AI Technical Summary
The existing photography posture assistance system has shortcomings in dynamic adaptability, interaction efficiency, posture generation stability and composition selection, resulting in poor user shooting experience.
Using a photography composition generation and photography prompt system based on the scene material library, intelligent photography posture guidance is provided, combining deep learning and expert experience to achieve real-time adaptation and dynamic shooting.
It improves the user's shooting experience, provides professional and accurate shooting guidance, ensures that the posture and scene are highly compatible, and captures the best moments dynamically, improving shooting quality and creative possibilities.
Smart Images

Figure CN120336566A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent photography technology, and in particular to a photography composition generation and photo-taking prompting system based on a scene material library. Background Art
[0002] Pictures play an irreplaceable and key role in today's social media. Smartphone photography has become the main way for users to take photos in daily life. However, ordinary users often have unsatisfactory photo quality due to stiff postures and poor composition. There are some existing photo posture assistance solutions in the prior art, such as generating a framework of photo postures for the photographer through preset classifications through set fixed templates, or using artificial intelligence to generate photo postures that match the current scene for users. However, these methods also have obvious limitations. The fixed template effect is rigid, the posture and the photo scene may not be coordinated, the composition may not be beautiful, the user needs to manually select the template, the interaction efficiency is low, the content generated by artificial intelligence is poorly stable, and overly conservative or exaggerated postures will appear, which makes it difficult for users to use them when taking photos. Although some studies have tried to solve the difficulty of posture recommendation, it is difficult for users to accurately fine-tune the details of the posture. Therefore, there is an urgent need for a photo posture assistance system that can adapt to the shooting environment in real time, provide visual comparison and intelligent prompts, so as to solve the problems of poor dynamic adaptability, low interaction efficiency and insufficient guidance accuracy of the prior art.
[0003] In summary, in today's social media era, although major social platforms that focus on content creation and sharing have launched some auxiliary photo functions, the existing functions are of limited practical help to users. The current mainstream photo assistance methods have the following obvious defects: 1) Difficulty in matching templates with scenes: Traditional preset fixed templates are seriously inflexible in practical applications. The shooting poses they provide are difficult to coordinate and match with complex and diverse real-life scenes. They cannot meet users’ rich and diverse personalized creation needs, which greatly limits users’ creative expression.
[0004] 2) Bottleneck of pose generation stability: Although AI-based pose generation technology is advanced to a certain extent, it is limited by the current technological level. The stability of the generated results is poor, and the generated poses are often overly conservative or exaggerated. It cannot provide users with reliable and practical references for photo poses, which seriously affects the user's shooting experience.
[0005] 3) Shortcomings in dynamic shooting and composition selection: Existing photo assistance functions mostly focus on static images, with limited ability to capture dynamic information during the shooting process. In addition, there is a lack of intelligent analysis and automatic screening mechanisms in composition selection, which makes it very easy for users to miss the best moments for shooting and it is difficult to obtain ideal shooting effects.
[0006] 4) Problems such as poor coordination between posture and scene, and insufficient generation stability. Summary of the Invention
[0007] The object of the present invention is to provide a photography composition generation and photo-taking prompt system based on a scene material library in view of the deficiencies of the prior art. By adopting a method based on a fused scene material library, a photography composition generation and photo-taking prompt system is constructed. Through real-time pose recognition, intelligent recommendations for portrait compositions that provide photo-taking pose guidance for users are provided, and a comprehensive solution integrating high-quality pose recommendations, precise photo-taking guidance, and intelligent dynamic shooting functions is provided for non-professional photography users and people lacking photo-taking experience, comprehensively improving the shooting experience and creation quality of users. The system combines a scene material library, expert experience, and innovative algorithms for photography composition generation and photo-taking prompts, covering static photo-taking assistance and LivePhoto dynamic shooting functions, and can provide real-time adaptation to the shooting environment, visual comparison, and intelligent prompts for photo-taking poses, effectively solving problems such as poor dynamic adaptability, low interaction efficiency, poor coordination between poses and scenes, insufficient generation stability, and insufficient guidance accuracy, and having good application prospects and commercial development value.
[0008] The object of the present invention is achieved as follows: A photography composition generation and photo-taking prompt system based on a scene material library, characterized in that the system is composed of a scene acquisition and feature extraction module, a template matching and retrieval module, a pose generation and optimization module, an aesthetic evaluation module, a real-time interaction prompt module, and a dynamic shooting and optimal frame selection module, to realize intelligent recommendations for portrait compositions with photo-taking pose guidance. The scene acquisition and feature extraction module will obtain real-time user shooting scene images, segment the portrait and the background, and extract the global semantic features, HSV color histograms, and local texture features of the background to construct a multi-dimensional feature vector; the template matching and retrieval module, based on the scene feature vector, retrieves the template with the highest matching degree with the current scene from the pre-constructed template human pixel material library, and sorts it comprehensively according to color similarity, texture similarity, and semantic similarity; the pose generation and optimization module extracts the skeletal key points of the portrait in the matching template through the OpenPose algorithm to generate skeleton pose guidance information, and combines the expert experience knowledge base to optimize the pose; the aesthetic evaluation module uses the CLIP model to extract the image semantic features, and combines the MLP network to score the candidate poses, and its evaluation follows the rule of thirds, color coordination, pose naturalness, and picture balance; the real-time interaction prompt module uses a generative AI model to dynamically generate pose adjustment suggestions and scene interaction prompts, and converts them into colloquial instructions for output through natural language processing technology; the dynamic shooting and optimal frame selection module continuously shoots multiple frames of images, and based on the composition score, screens the optimal frame in real time, and saves the static photo and the complete video sequence.
[0009] The specific implementation steps of the intelligent recommendation for portrait composition with photo-taking pose guidance include: 1) Extraction of high-quality portrait photos and construction of a material library The system platform stores a large amount of user photo data, and uses advanced classification and annotation technology to accurately mark portrait images. With the help of computer vision-based technical means, and integrating multi-dimensional data such as user comments and likes, high-quality photos are screened and extracted. In order to improve the efficiency of subsequent retrieval, these photos are carefully classified and stored according to scene types. The selected high-quality photos are retained separately as "templates", and a large and detailed template portrait material library is built on the back end of the application platform.
[0010] 2) Scene collection and perception When the user starts the photo function, the camera immediately captures the current photo background image, and at the same time uses advanced image recognition technologies such as deep learning-driven scene classification models on the mobile terminal to quickly and accurately identify and determine the category of the current scene.
[0011] 3) Scenario-based template retrieval The mobile terminal transmits the scene information obtained in real time to the platform system. The platform uses a carefully designed retrieval algorithm to efficiently search the template portrait material library and accurately find the high-quality "template" that matches the current scene most. The retrieval algorithm comprehensively considers multiple key features such as the color distribution, object layout, and environmental atmosphere of the scene to ensure the high accuracy of the retrieval results.
[0012] 4) Template pose optimization and reference generation The retrieved "template" is processed using the OpenPose algorithm to accurately extract the key point positions and skeleton structure information of the human body in the template. By analyzing this information and combining it with expert experience and knowledge, the template posture is optimized and adjusted to generate accurate posture guidance information for user reference.
[0013] 5) Expert experience and quantification of beautification effects With the help of knowledge engineering methods, the rich human photography experience is formally represented to build a professional experience knowledge base. The knowledge base is used to test the pose generation results to determine whether they are consistent with human actual photography experience. At the same time, a series of scientific and reasonable logical rules are set to quantitatively evaluate the pose generation quality from multiple dimensions such as the naturalness of the pose, coordination with the scene, and visual beauty.
[0014] 6) Interactive prompt information generation According to the optimized template poses, expert experience, and quantitative evaluation results, natural language processing technology is used to dynamically generate intuitive, easy-to-understand, and vivid prompt information for the photographer. These prompt information not only include specific guidance on pose adjustment but also cover relevant suggestions for interacting with the scene. The photographer can adjust the shooting pose and camera parameters in real time according to these prompt information to achieve the best shooting effect.
[0015] 7) Live Photo Dynamic Shooting The platform system automatically starts continuous shooting to capture a short video sequence of about 30 frames. For each frame image, a composition scoring system is used to conduct real-time evaluation based on five key indicators: the degree of compliance with the rule of thirds, sharpness, color coordination, subject prominence, and picture balance. The composition score is obtained by weighted summation. After shooting, the system automatically analyzes the scores of each frame and selects the frame with the highest score as a static photo for saving, and at the same time stores the complete video sequence.
[0016] The present invention has the following beneficial technical effects and significant technological progress compared with the prior art: 1) Simplification of operation process and optimization of experience: Innovatively realizes intelligent scene recognition and template generation technology, optimizing the cumbersome operation mode of manual template selection by users in traditional photography applications. Through powerful deep learning algorithms, the system can instantly and accurately capture the current environmental characteristics and automatically recommend the most suitable pose template for users, enabling users to fully immerse themselves in shooting creation, achieving seamless optimization of technology for the photography experience, and greatly improving the convenience of user operation and the concentration during the shooting process.
[0017] 2) Enhancement of aesthetic value and scene matching degree: Pioneeringly integrates an aesthetic scoring model and multi-dimensional feature real-time matching technology. Starting from multiple key dimensions such as composition, color, and pose, it intelligently evaluates the aesthetic value of the template and its matching degree with a specific scene. Not only ensures that the generated content has excellent visual beauty but also ensures a high degree of fit between the template and the actual scene, providing professional, accurate, and elegant shooting guidance for users and significantly improving the artistic quality of shooting works.
[0018] 3) Real-time interaction and expansion of shooting ideas: Develops a real-time interaction system with powerful semantic understanding ability, different from the common static and single-pose guidance modes in the market. Through natural language processing technology, the system can accurately provide users with dynamic and vivid pose guidance and scene interaction suggestions. At the same time, by means of relevant content of text prompts and scene interaction, it stimulates users' new shooting ideas and creativity, greatly enriching the user shooting experience and creative possibilities.
[0019] 4) Solving User Pain Points and Creating Market Value: It precisely addresses the core pain points faced by novice photographers in aspects such as posing and selecting appropriate angles. Through multi-dimensional intelligent perception, scene-based personalized recommendations, and real-time dynamic adjustment mechanisms, it comprehensively enhances the user's shooting experience. The intelligent recommendation algorithm based on the scene material library not only has broad application prospects at the individual user level but also can explore huge market value through in-depth cooperation with social platforms, photography equipment manufacturers, and content creation agencies. In addition, this technology can be widely extended to many fields such as e-commerce, advertising, and traditional commerce, providing intelligent and customized content creation support for various industries and strongly promoting the digital and intelligent transformation process of society.
[0020] 5) Advantages of Dynamic Shooting and Intelligent Composition: It realizes the efficient capture of dynamic images and the selection of intelligent composition, effectively preventing users from missing perfect shooting moments. The system automatically selects the frame with the best composition from continuous frames as a static photo while completely saving the video sequence, providing the dynamic content before and after the static photo and enriching the form of shooting records. The intelligent composition evaluation mechanism combined with professional photography composition rules further improves the quality and artistic value of shooting works, meeting the user's needs for diverse and high-quality shooting functions. Brief Description of the Drawings
[0021] Figure 1 It is a schematic structural diagram of the present invention; Figure 2 It is a schematic diagram of the platform system of Embodiment 1; Figure 3 It is a schematic diagram of the dynamic template matching effect; Figure 4 It is a schematic diagram of aesthetic score and multi-dimensional evaluation; Figure 5 It is a schematic diagram of generating real-time shooting prompt information based on mobile perception; Figure 6 It is a schematic diagram of live Photo dynamic shooting and saving the optimal frame. Detailed Embodiment
[0022] Refer to Figure 1 , the overall architecture of the present invention is divided into four core modules, namely, the scene recognition and feature extraction module, the template matching and retrieval module, the pose generation and aesthetic evaluation module, the real-time interaction and prompt generation module, and the dynamic shooting and optimal frame selection module.
[0023] When the user starts the shooting function in the scene recognition and feature extraction module, the camera will capture the scene image in real time. The system will use DeepLabV3 to segment the human figure and the background, and at the same time, construct a feature vector in multiple dimensions such as combining the global features of the cut background, the HSV color histogram, and the local texture features.
[0024] The template matching and retrieval module will comprehensively consider various aspects such as color, texture, and semantics for the extracted features, and strive to find the most suitable template from a high-quality scene material library. Among them, the color similarity is measured by calculating the cosine similarity of the HSV histogram; the texture similarity measures the difference in local texture distribution through the EMD algorithm; the semantic similarity is sorted based on the L2 distance of ResNet50 features.
[0025] The pose generation and aesthetics evaluation uses OpenPose to detect the key points of the skeleton of the optimal template retrieved (18 key points), and generates a skeleton diagram for the user to refer to in order to achieve a finished film effect that meets the user's expectations. At the same time, the clip model is used to extract the features of the user's photo, and the model is trained by comprehensively considering multiple dimensions such as composition, color, and pose, so that it can perform aesthetic scoring on the candidate poses to intelligently evaluate the aesthetic value and scene matching degree of the template.
[0026] The real-time interaction and prompt generation module uses an AI model deployed on the intranet. According to the user's pose, scene features, and aesthetic scoring results, it dynamically generates structured instructions (in JSON format), converts the JSON instructions into colloquial prompts, and returns them to the front-end page for the user to refer to. In the generation of pose adjustment suggestions, the system will real-time identify the visible parts of the user's body based on the 18 key point detections of OpenPose (such as whether the arm is blocked) to ensure the beauty of the pose. The real-time interaction includes: pose adjustment suggestions and scene interactions that provide the user with new shooting ideas.
[0027] The dynamic shooting and optimal frame selection module realizes dynamic shooting and best frame selection through continuous multi-frame shooting + intelligent scoring and screening. The mobile camera is called to continuously shoot at a rate of 30 frames per second, generating a short video sequence (30 frames) with a duration of about 1 second, and each frame image is evaluated in terms of five key indicators including the degree of compliance with the rule of thirds, sharpness, color coordination, subject prominence, and picture balance, as well as the collected expert experience, and reasonable weights are assigned to it for multi-dimensional evaluation. The system will save the optimal frame, and at the same time, in order to meet the user's personal preferences, the system will save its complete video sequence for the user to choose.
[0028] The present invention will be further described below in conjunction with specific embodiments. Embodiment
[0029] Refer to Figure 2 , and perform intelligent recommendation for portrait composition according to the following steps: Step ①: Extract high-quality photos based on the existing technology and save them separately as "templates", and build a template human pixel material library on the back end of the application platform.
[0030] Step ②: The user's captured scene is input through the camera. The system uses DeepLabV3 to separate the human figure from the background and retain the scene features.
[0031] Step ③: ResNet50 extracts global semantic features, constructs a multi-dimensional feature vector by combining the HSV color histogram and local texture features, and matches the optimal template from the scene material library based on the L2 distance, cosine similarity, and EMD algorithm.
[0032] Step ④: OpenPose detects the key points of the human figure pose in the matched template and generates a skeleton pose. The user will improve the quality and effect of the finished video by imitating the key points of the pose in the high-quality matched template.
[0033] Step ⑤: Generate interactive information prompts through natural language processing. At the same time, use the CLIP+MLP model to perform aesthetic scoring on the candidate poses and overlay it on the user interface.
[0034] Step ⑥: Collect expert experience and quantify the beautification effect Step ⑦: Support dynamic shooting. Based on expert experience, perform real-time evaluation through multi-dimensional indicators to obtain the optimal live frame and save the complete live video at the same time.
[0035] Refer to Figure 3 a, The real-time scene and template matching overlay effect of using the CLIP+MLP model to perform aesthetic scoring on the candidate poses and overlay it on the user interface.
[0036] Refer to Figure 3 b, High-quality templates and pose recommendations obtained by using the CLIP+MLP model to perform aesthetic scoring and matching on the candidate poses.
[0037] Refer to Figure 4 , Real-time extract the image semantic features through the CLIP model, and use the MLP network to comprehensively output scores in multiple dimensions such as composition, color coordination, and pose naturalness.
[0038] Refer to Figure 5 , Through natural language processing technology, the system can accurately provide users with dynamic and vivid pose guidance and scene interaction suggestions (as shown on the right side of the figure).
[0039] Refer to Figure 6 , The system automatically analyzes the scores of each frame of a short video sequence of about 30 frames, and selects the frame with the highest score to be saved as a static photo.
[0040] The above embodiments are only further explanations of the present invention, and are not intended to limit this patent. All equivalent implementations of the present invention should be included within the scope of the claims of this patent.
Claims
1. A photography composition generation and photo-taking prompt system based on a scene material library, characterized in that A photography composition generation and photo-taking prompt system composed of a scene acquisition and feature extraction module, a template matching and retrieval module, a pose generation and optimization module, an aesthetic evaluation module, a real-time interaction prompt module, and a dynamic shooting and optimal frame selection module realizes intelligent recommendations for photo-taking pose guidance and portrait composition. The scene acquisition and feature extraction module will obtain real-time user shooting scene images, segment the portrait from the background, and extract the global semantic features, HSV color histograms, and local texture features of the background to construct a multi-dimensional feature vector. The template matching and retrieval module, based on the scene feature vector, retrieves the template with the highest matching degree with the current scene from the pre-constructed template human pixel material library, and sorts it comprehensively according to color similarity, texture similarity, and semantic similarity. The pose generation and optimization module extracts the skeletal key points of the portrait in the matching template through the OpenPose algorithm to generate skeleton pose guidance information, and combines the expert experience knowledge base for pose optimization. The aesthetic evaluation module uses the CLIP model to extract the image semantic features, and combines the MLP network to score the composition of the candidate poses. Its evaluation of the rule of thirds compliance, color coordination, pose naturalness, and picture balance. The real-time interaction prompt module uses the generative AI model to dynamically generate pose adjustment suggestions and scene interaction prompts, and converts them into colloquial instructions through natural language processing technology for output. The dynamic shooting and optimal frame selection module continuously shoots multiple frames of images, and based on the composition score, it real-time filters the optimal frame, and saves the static photo and the complete video sequence.
2. The photography composition generation and photo-taking prompt system based on a scenario material library according to claim 1, wherein The specific steps of the intelligent recommendation for photo-taking pose guidance and portrait composition include: (1) Construction of the template human pixel material library Using the user photo data stored in the photography composition generation and photo-taking prompt system, adopt classification technology to mark portrait pictures, and based on computer vision technology, fuse multi-dimensional data such as user comments and likes, and screen and extract portrait scene photos from them as templates to construct a template human pixel material library; (2) Scene acquisition and perception Use a scene classification model driven by deep learning to identify the collected images and determine the category of the current scene; (3) Template retrieval based on the scene Input the identified scene information into the photography composition generation and photo-taking prompt system in real time, and use a retrieval algorithm to retrieve in the template human pixel material library, including: key features such as the color distribution of the scene, object layout, and environmental atmosphere, and find the template with the highest matching degree with the current scene; (4) Template pose optimization and reference generation Use the OpenPose algorithm to extract the key point positions and skeleton structure information of the human body from the retrieved template, and optimize and adjust the template pose according to expert experience knowledge to generate accurate pose guidance information for users to reference; (5) Expert experience and quantification of beautification effects Formalize the human photo-taking experience using knowledge engineering methods to construct an experience knowledge base, and use this knowledge base to detect the pose generation results to judge whether they conform to human actual photo-taking experience, and quantitatively evaluate the quality of pose generation from the aspects of pose naturalness, coordination with the scene, and visual beauty; (6) Generation of interactive prompt information According to the optimized template poses, expert experience, and quantitative evaluation results, natural language processing technology is used to dynamically generate prompt information for the photographer. The photographer adjusts the shooting pose and camera parameters in real time according to the prompt information to achieve the best shooting effect. The prompt information includes: specific guidance on pose adjustment, as well as relevant suggestions covering interactions with the scene; (7) Live Photo Dynamic Shooting The photography composition generation and photo-taking prompt system starts continuous shooting, captures a short video sequence of about 30 frames, uses the aesthetic evaluation module to perform real-time evaluation on each frame of the image, calculates the composition score by weighted summation, selects the frame with the highest score as a static photo for storage, and stores the complete video sequence at the same time. The composition score is evaluated in real time based on the degree of compliance with the rule of thirds, sharpness, color coordination, subject prominence, and picture balance.
Citation Information
Cited By
Photography auxiliary method, system and equipment based on artificial intelligence and medium
CN121567952A