Photographing control system based on pre-prepared scene template and structured parameter setting driving
By establishing a global spatial benchmark, feature calibration, and temporal synchronization, a shooting semantic parsing model is constructed to generate hierarchical control logic instructions. This solves the problems of data fragmentation and insufficient autonomous semantic parsing capabilities in existing shooting control technologies, and achieves high-precision, adaptive shooting control effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI MOGONG CULTURE COMM CO LTD
- Filing Date
- 2026-03-31
- Publication Date
- 2026-07-03
AI Technical Summary
Existing shooting control technologies struggle to achieve high-precision, adaptive, and standardized shooting control, exhibiting issues such as data fragmentation, lack of autonomous semantic parsing capabilities, command conflicts, asynchronous execution, and poor scene versatility, thus failing to meet the demands of intelligent shooting.
By establishing a unified global spatial benchmark, performing feature space coordinate calibration and temporal synchronization, constructing a full dataset of images, building a five-layer architecture for image semantic parsing, performing subject motion temporal analysis, generating hierarchical control logic commands, and comparing the image shooting effects in real time, the commands are dynamically corrected to optimize the actions of the shooting device.
It improves the accuracy of shooting control data and enhances strategy adaptability, avoids command conflicts and asynchronous actions, improves the synergy of shooting effects and scene adaptability, and achieves stability and accuracy in the shooting process.
Smart Images

Figure CN122340348A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent shooting control technology, specifically involving a shooting control system driven by pre-made scene templates and structured parameter settings. Background Technology
[0002] Currently, filming operations in scenarios such as film and television recording, live shooting, and product image acquisition are gradually developing towards automation and intelligence. However, existing shooting control technologies still have many technical shortcomings, making it difficult to meet the requirements of high-precision, adaptive, and standardized shooting control. Existing shooting control systems mostly rely on manual operation of shooting devices, requiring operators to adjust camera positions, gimbal attitudes, and lens parameters in real time. The operation process is cumbersome, the work efficiency is low, and the shooting effect is highly dependent on the operator's experience, resulting in poor consistency of shooting results for different scenes and different personnel.
[0003] Some automated shooting systems use only a single sensing module to collect data, which cannot simultaneously acquire full-dimensional information on the shooting scene space, the movement of the subject being shot, and the shooting execution status. They have not established a unified global spatial benchmark, and the feature data suffers from spatiotemporal asynchrony and data fragmentation, failing to provide complete data support for shooting control. At the same time, existing systems lack the ability to semantically analyze the movement of the subject, and can only execute fixed preset instructions. They cannot autonomously match the shooting logic according to the subject's movement behavior, resulting in extremely poor adaptability to dynamic subjects.
[0004] In addition, the control commands generated by the existing system lack hierarchical logic and compliance verification. The execution components of the shooting device, such as camera position adjustment, gimbal rotation, and lens adjustment, are difficult to coordinate, which can easily lead to command conflicts and asynchronous execution. Most systems have not built an effect comparison and deviation correction mechanism, so they cannot calculate the shooting effect deviation in real time and iteratively optimize the control commands, making it difficult to guarantee the stability and accuracy of the shooting process.
[0005] Existing template-based shooting control solutions use fixed parameter templates, which cannot be dynamically adapted and adjusted according to on-site conditions. They have weak scene versatility and cannot adapt to complex and ever-changing shooting environments and diverse shooting needs, thus restricting the large-scale application of intelligent shooting control technology. Summary of the Invention
[0006] To address the aforementioned problems in the existing technology, this invention provides a shooting control system based on pre-made scene templates and structured parameter settings. The objective of this invention can be achieved through the following technical solutions: include: The spatial perception unit establishes a unified global spatial benchmark based on the acquired full-dimensional features of the images; it performs spatial coordinate calibration and temporal synchronization processing on various features, and classifies and integrates the calibrated and synchronized features to form a full dataset of the images. The semantic modeling unit, based on the full dataset of the captured images and combined with pre-made scene parameter templates, constructs a shooting semantic parsing model; performs temporal analysis and behavior parsing on the motion characteristics of the subject being shot, obtains the shooting semantics corresponding to the subject's motion, and generates corresponding shooting control strategies; The parameter scheduling unit performs a full-dimensional breakdown and logic compliance verification of the shooting control strategy; and generates directly executable hierarchical control logic instructions based on the verified strategy according to the shooting device control logic. The adaptive execution unit performs structural adjustments to the shooting device and controls the entire shooting process according to the control logic instructions; it compares the actual effect characteristics of the real-time acquired images with the preset target requirements in all dimensions; it calculates the deviation value between the actual effect characteristics and the target requirements, dynamically corrects instructions according to the degree of deviation value, and iteratively optimizes the actions and related parameters of the shooting device.
[0007] Specifically, the process of establishing a unified global spatial benchmark is as follows: Obtain the physical spatial layout of the shooting scene, select a fixed point in the scene as the origin of the coordinate system, and preset the direction of the coordinate axis extension and the coordinate range covering the effective shooting area. Selected reference points are spatially located and marked, and a three-dimensional spatial coordinate system for the shooting scene is constructed based on the marked points; The three-dimensional spatial coordinate system is adjusted based on the shooting conditions to match the range of motion of the subject being shot with the operating range of the shooting device.
[0008] Specifically, the process of spatial coordinate calibration and temporal synchronization of various features is as follows: Extract the original spatial information during the collection of various features, map the spatial data of various features to the coordinate system of the global spatial reference, correct the spatial position deviations caused by various features, and assign a unified spatial coordinate identifier to various features; Extract the collection time information of various features, align the timelines of different feature data with a unified timeline as the benchmark, fill in the missing time nodes of different feature data, and remove redundant and invalid data on the timeline.
[0009] Specifically, the process of forming the full dataset is as follows: Distinguish between the spatial characteristics of the shooting scene, the motion characteristics of the subject being shot, and the execution state characteristics of the shooting control; Distinguish between each type of core feature and standardize the format of different feature data; Obtain the data storage format and dimensional relationship of various features, and establish mutual retrieval relationships for different feature data.
[0010] Specifically, the semantic parsing model includes a template parsing layer, an architecture determination layer, a data import layer, a model training layer, and a logic optimization layer. The specific construction process is as follows: The input consists of the full dataset of images captured and pre-made scene parameter templates; the output is the image semantic parsing model. The template parsing layer retrieves pre-made scene parameter templates and parses the preset shooting parsing dimensions, feature analysis logic, and shooting semantic association rules; the architecture determination layer determines the overall architecture and core output direction based on the parsing results; the data import layer imports the full shooting dataset and integrates the feature association relationship between the shooting scene and the shooting subject; the model training layer continuously trains the parsing capability in combination with the feature association relationship; and the logic optimization layer optimizes the feature matching logic.
[0011] Specifically, the process of performing temporal analysis and behavioral analysis on the motion characteristics of the subject being photographed is as follows: Extract motion feature data of the subject being photographed, remove invalid interference information from the data, and form continuous time series data of the subject's motion; The analysis focuses on the direction, amplitude, and rhythm of the subject's movement at different time stages. Based on the semantic association logic of the scene parameter template, the shooting behavior corresponding to each motion state of the subject is analyzed to obtain the core shooting semantics corresponding to the overall motion of the subject.
[0012] Specifically, the process of performing full-dimensional decomposition and logical compliance verification is as follows: The shooting control strategy is broken down into specific control requirements for scene adaptation control, subject tracking control, camera movement trajectory control, and lens parameter control. To obtain the actual execution capabilities of the shooting device, and based on the actual spatial conditions of the shooting scene, to check each of the specific control requirements after disassembly, and to eliminate specific control requirements that do not match the equipment capabilities and scene conditions.
[0013] Specifically, the process of generating directly executable hierarchical control logic instructions is as follows: The hierarchical division and dedicated control logic of each component of the preset shooting device; The verified control requirements for each dimension are allocated according to the execution component level and converted into control information that can be recognized by the corresponding execution component. The main control instructions and auxiliary control instructions are divided according to the importance of the control requirements and the execution order, and then integrated to form a hierarchical control logic instruction.
[0014] Specifically, the process of structural adjustment of the shooting device and full-process control of the shooting flow is as follows: Perform hierarchical parsing of control logic instructions to obtain the adjustment requirements and action ranges corresponding to each execution component; The actuators of the shooting device are driven to perform corresponding adjustment actions, and the adjustment status of each actuator is monitored in real time. Based on the preset shooting rhythm and execution sequence, control the entire process of starting and stopping shooting, switching shot sizes, and adjusting camera movements.
[0015] Specifically, the process of comparing the real-time acquired image effect features with the preset target requirements in all dimensions is as follows: Real-time capture of the captured footage, and extraction of actual effect features from the real-time captured footage; For each actual effect feature, quantitative extraction and corresponding characterization are performed. Based on the preset shooting target requirements, the corresponding judgment criteria and quantitative indicators for each dimension are obtained. The quantitative data of actual features are compared with the corresponding judgment criteria, and the comparison results and differences are recorded.
[0016] Specifically, the process of calculating the deviation between the actual effect characteristics and the target requirements is as follows: For comparison dimensions where there are differences, the degree of difference between the actual effect characteristics and the target requirements is quantitatively calculated according to the preset unified evaluation standard. Based on the degree of influence of each dimension on the overall shooting, a weight ratio is set for different comparison dimensions; The weighted calculation of the degree of difference in each dimension is performed, and the weighted difference values of all dimensions are integrated and summarized. A comprehensive deviation value of the overall shooting effect is formed according to a unified calculation rule, and the specific deviation value of each single dimension is extracted and retained.
[0017] Specifically, the process of dynamically correcting the instructions and iteratively optimizing the actions and related parameters of the shooting device is as follows: The preset deviation value classification standard determines the deviation level of the overall shooting effect; Develop corresponding instruction correction schemes for different deviation levels, and locate the cause of shooting deviation and the corresponding control dimension based on the single-dimensional deviation value; Adjust the relevant parameters of the corresponding dimension in the control logic command, send the corrected control logic command to the shooting device, and iteratively adjust the actions and parameters of the shooting device.
[0018] The beneficial effects of this invention are as follows: (1) By setting up a technical structure that includes global spatial benchmark construction, feature space coordinate calibration and temporal synchronization, full dataset integration, five-layer architecture shooting semantic analysis model construction, and subject motion temporal analysis and behavior analysis, the spatiotemporal calibration and classification integration of shooting full-dimensional features can be completed in a standardized manner, forming complete and standardized data support. At the same time, the shooting semantics corresponding to the subject motion can be accurately analyzed, solving the problems of fragmented traditional shooting data and lack of autonomous semantic analysis capabilities, and improving the data accuracy and strategy adaptability of shooting control. (2) By setting up a series of execution processes including full-dimensional decomposition and logical compliance verification of shooting control strategy, generation of hierarchical control logic instructions, adjustment of shooting device structure and process control, comparison of shooting effect features, quantitative calculation of deviation value, dynamic correction of control instructions and iterative optimization of parameters, the problems of shooting device instruction conflict and asynchronous action can be avoided. At the same time, the closed-loop control of shooting effect can be realized, which greatly improves the synergy, accuracy and scene adaptability of shooting control. Attached Figure Description
[0019] To facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings.
[0020] Figure 1 This is a system architecture diagram of the shooting control system based on pre-made scene templates and structured parameter settings of the present invention; Figure 2 This is a data flow diagram of the shooting control system based on pre-made scene templates and structured parameter settings of the present invention. Detailed Implementation
[0021] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided.
[0022] Please see Figure 1-2 A shooting control system driven by pre-made scene templates and structured parameter settings; include: The spatial perception unit establishes a unified global spatial benchmark based on the acquired full-dimensional features of the images; it performs spatial coordinate calibration and temporal synchronization processing on various features, and classifies and integrates the calibrated and synchronized features to form a full dataset of the images. The semantic modeling unit, based on the full dataset of the captured images and combined with pre-made scene parameter templates, constructs a shooting semantic parsing model; performs temporal analysis and behavior parsing on the motion characteristics of the subject being shot, obtains the shooting semantics corresponding to the subject's motion, and generates corresponding shooting control strategies; The parameter scheduling unit performs a full-dimensional breakdown and logic compliance verification of the shooting control strategy; and generates directly executable hierarchical control logic instructions based on the verified strategy according to the shooting device control logic. The adaptive execution unit performs structural adjustments to the shooting device and controls the entire shooting process according to the control logic instructions; it compares the actual effect characteristics of the real-time acquired images with the preset target requirements in all dimensions; it calculates the deviation value between the actual effect characteristics and the target requirements, dynamically corrects instructions according to the degree of deviation value, and iteratively optimizes the actions and related parameters of the shooting device.
[0023] In this embodiment, the full-dimensional shooting features refer to the three types of core features collected and processed by the system, specifically including the spatial features of the shooting scene, the motion features of the shooting subject, and the execution state features of the shooting control. These are the basic data sources for the system to carry out spatial benchmark construction, data integration, and semantic analysis. Pre-built scene parameter templates refer to standardized scene parameter files that are pre-compiled and stored by the system. They contain built-in parameter benchmarks, parsing rules, control logic and semantic association conditions for different shooting scenarios, which can be directly called to adapt to the control requirements of the corresponding shooting scenarios. Structured parameters refer to standardized parameters organized according to a unified format, hierarchical relationship and correlation. Specifically, they include scene space parameters, subject motion parameters and shooting control parameters, and have the characteristics of being decomposable, verifiable, convertible and schedulable. A unified global spatial reference refers to a three-dimensional spatial coordinate system built based on fixed reference points of the shooting scene, including the coordinate origin, coordinate axis direction and effective shooting coverage, used to unify the spatial positioning of all features and match the movement range of the shooting subject with the operating range of the shooting device. The preset shooting analysis dimensions refer to the semantic analysis classification dimensions pre-defined in the scene parameter template, specifically including scene adaptation analysis dimension, subject tracking analysis dimension, camera movement trajectory analysis dimension, and lens parameter analysis dimension, which are the classification basis for the model to carry out semantic analysis. Feature analysis logic and shooting semantic association rules refer to the preset technical rules in the scene parameter template. Feature analysis logic is the analysis method for the system to process spatial, motion, and execution state features; shooting semantic association rules are the corresponding matching rules between feature data and shooting semantics. The semantic association logic of the scene parameter template refers to the derivation logic built into the scene parameter template, which is used to establish the corresponding derivation relationship between the motion characteristics of the shooting subject, the spatial characteristics of the shooting scene, and the shooting semantics and shooting control strategies. Shooting action direction refers to the specific shooting execution direction corresponding to different movement states of the subject, including shooting execution directions such as camera position adjustment, gimbal attitude adjustment, lens parameter change, framing switch, and camera movement trajectory planning; Core shooting semantics refers to the core shooting control requirements extracted from the overall movement of the subject. It is the core basis for generating shooting control strategies and covers core shooting control semantics such as subject tracking, camera movement control, lens adjustment, and scene adaptation.
[0024] Specifically, the process of establishing a unified global spatial benchmark is as follows: Obtain the physical spatial layout of the shooting scene, select a fixed point in the scene as the origin of the coordinate system, and preset the direction of the coordinate axis extension and the coordinate range covering the effective shooting area. Selected reference points are spatially located and marked, and a three-dimensional spatial coordinate system for the shooting scene is constructed based on the marked points; The three-dimensional spatial coordinate system is adjusted based on the shooting conditions to match the range of motion of the subject being shot with the operating range of the shooting device.
[0025] Specifically, the process of spatial coordinate calibration and temporal synchronization of various features is as follows: Extract the original spatial information during the collection of various features, map the spatial data of various features to the coordinate system of the global spatial reference, correct the spatial position deviations caused by various features, and assign a unified spatial coordinate identifier to various features; Extract the collection time information of various features, align the timelines of different feature data with a unified timeline as the benchmark, fill in the missing time nodes of different feature data, and remove redundant and invalid data on the timeline.
[0026] Specifically, the process of forming the full dataset is as follows: Distinguish between the spatial characteristics of the shooting scene, the motion characteristics of the subject being shot, and the execution state characteristics of the shooting control; Distinguish between each type of core feature and standardize the format of different feature data; Obtain the data storage format and dimensional relationship of various features, and establish mutual retrieval relationships for different feature data.
[0027] Specifically, the semantic parsing model includes a template parsing layer, an architecture determination layer, a data import layer, a model training layer, and a logic optimization layer. The specific construction process is as follows: The input consists of the full dataset of images captured and pre-made scene parameter templates; the output is the image semantic parsing model. The template parsing layer retrieves pre-made scene parameter templates and parses the preset shooting parsing dimensions, feature analysis logic, and shooting semantic association rules; the architecture determination layer determines the overall architecture and core output direction based on the parsing results; the data import layer imports the full shooting dataset and integrates the feature association relationship between the shooting scene and the shooting subject; the model training layer continuously trains the parsing capability in combination with the feature association relationship; and the logic optimization layer optimizes the feature matching logic.
[0028] Specifically, the process of performing temporal analysis and behavioral analysis on the motion characteristics of the subject being photographed is as follows: Extract motion feature data of the subject being photographed, remove invalid interference information from the data, and form continuous time series data of the subject's motion; The analysis focuses on the direction, amplitude, and rhythm of the subject's movement at different time stages. Based on the semantic association logic of the scene parameter template, the shooting behavior corresponding to each motion state of the subject is analyzed to obtain the core shooting semantics corresponding to the overall motion of the subject.
[0029] Specifically, the process of performing full-dimensional decomposition and logical compliance verification is as follows: The shooting control strategy is broken down into specific control requirements for scene adaptation control, subject tracking control, camera movement trajectory control, and lens parameter control. To obtain the actual execution capabilities of the shooting device, and based on the actual spatial conditions of the shooting scene, to check each of the specific control requirements after disassembly, and to eliminate specific control requirements that do not match the equipment capabilities and scene conditions.
[0030] Specifically, the process of generating directly executable hierarchical control logic instructions is as follows: The hierarchical division and dedicated control logic of each component of the preset shooting device; The verified control requirements for each dimension are allocated according to the execution component level and converted into control information that can be recognized by the corresponding execution component. The main control instructions and auxiliary control instructions are divided according to the importance of the control requirements and the execution order, and then integrated to form a hierarchical control logic instruction.
[0031] Specifically, the process of structural adjustment of the shooting device and full-process control of the shooting flow is as follows: Perform hierarchical parsing of control logic instructions to obtain the adjustment requirements and action ranges corresponding to each execution component; The actuators of the shooting device are driven to perform corresponding adjustment actions, and the adjustment status of each actuator is monitored in real time. Based on the preset shooting rhythm and execution sequence, control the entire process of starting and stopping shooting, switching shot sizes, and adjusting camera movements.
[0032] Specifically, the process of comparing the real-time acquired image effect features with the preset target requirements in all dimensions is as follows: Real-time capture of the captured footage, and extraction of actual effect features from the real-time captured footage; For each actual effect feature, quantitative extraction and corresponding characterization are performed. Based on the preset shooting target requirements, the corresponding judgment criteria and quantitative indicators for each dimension are obtained. The quantitative data of actual features are compared with the corresponding judgment criteria, and the comparison results and differences are recorded.
[0033] Specifically, the process of calculating the deviation between the actual effect characteristics and the target requirements is as follows: For comparison dimensions where there are differences, the degree of difference between the actual effect characteristics and the target requirements is quantitatively calculated according to the preset unified evaluation standard. Based on the degree of influence of each dimension on the overall shooting, a weight ratio is set for different comparison dimensions; The weighted calculation of the degree of difference in each dimension is performed, and the weighted difference values of all dimensions are integrated and summarized. A comprehensive deviation value of the overall shooting effect is formed according to a unified calculation rule, and the specific deviation value of each single dimension is extracted and retained.
[0034] Specifically, the process of dynamically correcting the instructions and iteratively optimizing the actions and related parameters of the shooting device is as follows: The preset deviation value classification standard determines the deviation level of the overall shooting effect; Develop corresponding instruction correction schemes for different deviation levels, and locate the cause of shooting deviation and the corresponding control dimension based on the single-dimensional deviation value; Adjust the relevant parameters of the corresponding dimension in the control logic command, send the corrected control logic command to the shooting device, and iteratively adjust the actions and parameters of the shooting device.
[0035] In this embodiment, when establishing a unified global spatial reference, a fixed installation structure adapted to the shooting equipment is used as the basis. The top of a cylindrical shooting equipment with a diameter ranging from 70mm to 90mm is selected as the reference positioning carrier. The 120mm×80mm×40mm working area corresponding to the monitoring device is included in the coverage of the spatial reference. A fixed point within the shooting scene without displacement is selected as the origin of the coordinates. The extension direction of the coordinate axis is determined according to the installation position relationship between the shooting equipment and the monitoring device. The reference points within the scene are marked to establish a three-dimensional spatial coordinate system adapted to the shooting operation. The coordinate system parameters are adjusted in combination with the movement range of the shooting subject and the actual working range of the shooting device so that the spatial reference can completely cover the acquisition and processing range of all-dimensional features of the shooting.
[0036] In this embodiment, when performing spatial coordinate calibration and temporal synchronization processing on various features, the original spatial information of various features is first mapped to the established global spatial reference to correct the spatial position deviation of feature data caused by differences in acquisition location, and all features are given a unified spatial coordinate identifier. The temporal synchronization processing adopts hardware clock synchronization, using the system's unified crystal oscillator as the time reference, to synchronously acquire the captured full-dimensional feature data. The acquisition frequency is set to 10Hz, and the time error of feature data from different acquisition nodes is controlled within 1ms. At the same time, in accordance with the data transmission specification of the RS-485 communication interface, the missing time nodes of feature data are supplemented, and redundant and invalid data in the time dimension are eliminated to ensure the spatiotemporal consistency of all feature data.
[0037] In this embodiment, when constructing the shooting semantic parsing model, the STM32F103C8T6 microprocessor is used as the core processing unit. This processor is equipped with a 72MHz ARM Cortex-M3 core. The total input of the model is the full shooting dataset and the pre-made scene parameter templates. The total output is the shooting semantic parsing model that can complete the analysis of the subject's motion features. The model is divided into five layers: template parsing layer, architecture determination layer, data import layer, model training layer, and logic optimization layer. The template parsing layer calls up parameter templates that integrate ASC, GB / T 15406-2017, and SMPTE film and television lighting standards. The data import layer imports the full shooting dataset into the model architecture and incorporates the feature association relationship between the scene and the subject. Temperature compensation and linearization processing techniques are used during model training and optimization. The overall response time of the model is controlled within 50ms. The parsed data is stored in the onboard 1GB flash memory module.
[0038] In this embodiment, when performing a comprehensive breakdown and logical compliance verification of the shooting control strategy, the shooting control strategy is broken down layer by layer into four specific control requirements: scene adaptation control, subject tracking control, camera movement trajectory control, and lens parameter control. The verification process follows clear rules and standards. The physical limit verification of the equipment is based on the rated parameters of the shooting device's camera position movement range, gimbal rotation angle, and lens zoom. The kinematic constraint verification checks dynamic parameters such as the device's movement speed and rotation angular velocity. The scene safety verification ensures that the control commands will not cause the device to exceed the scene space boundary or collide. At the same time, combined with the spatial layout conditions of the shooting scene, each of the decomposed control requirements is checked one by one, and control content that does not match the shooting device's execution capabilities and scene space conditions is eliminated. The verification process follows the device communication protocol specifications such as ONVIF, VISCA, DMX512, and DALI to ensure that the decomposed control requirements can be transmitted and executed through standardized interfaces.
[0039] In this embodiment, when calculating the deviation between the actual effect characteristics and the preset target requirements, the deviation value adopts a definition method that combines multi-dimensional vectors and a single scalar. First, the independent deviations of multiple dimensions such as subject illumination, light ratio, color temperature consistency, image composition, and focus effect are calculated separately to form a multi-dimensional deviation vector. The deviation of each dimension is calculated by subtracting the measured value from the target value. Then, a weighting coefficient is set according to the degree of influence of each dimension on the shooting quality. The multi-dimensional deviation vector is integrated into a single comprehensive deviation scalar through a weighted summation fusion rule. The deviation calculation is based on the system's ±5% measurement accuracy to fully reflect the overall difference between the shooting effect and the target requirements.
[0040] In this embodiment, when dynamically correcting commands and iteratively optimizing the actions and related parameters of the shooting device based on the degree of deviation, a deviation level classification standard is preset. The trigger condition for adaptive execution is that the comprehensive deviation scalar exceeds the allowable range of ±5%. A three-level deviation level judgment standard is preset. The adjustment algorithm adopts a combination of hierarchical priority adjustment and closed-loop iterative correction. The control command parameters are gradually corrected according to the priority of subject illuminance, light ratio relationship, and color temperature consistency. After the correction command is issued, wait for 0.5 to 2 seconds for the adjustment to take effect, re-acquire the actual effect characteristics of the shooting image and calculate the deviation. The number of iterative corrections is controlled to be 3 to 5 times. The iterative optimization termination condition is that the comprehensive deviation scalar converges to within the target standard range of ±5%. After the termination condition is met, the actions and parameters of the shooting device are locked, and the shooting status is continuously monitored at a frequency of 10Hz.
[0041] In this embodiment, indoor fixed-position portrait photography is selected as the application scenario. The subject of the photography is a person in a fixed area within the scene. The shooting device is a fixed camera equipped with this system. The system operates based on preset portrait photography parameters. The specific process is as follows: After the system is powered on and initialized, the spatial perception unit enters the working state and continuously collects the full-dimensional shooting features S in the shooting scene. Features S cover three core contents: scene spatial distribution features, human movement features, and shooting device execution state features. The spatial perception unit uses fixed structures within the scene that do not change position as reference points to establish a unified global spatial reference M. Reference M completely covers the area where people are active and the area where the shooting device is operating, providing a unified spatial positioning basis for all feature data. Subsequently, the unit performs spatial coordinate calibration and temporal synchronization processing on the collected features S, mapping all feature data to the reference M to correct spatial position deviations. At the same time, it completes temporal alignment based on a unified clock, fills in missing time nodes, removes redundant and invalid data, and then classifies and organizes the calibrated and synchronized features according to attributes to form a structured full-scale shooting dataset D. The semantic modeling unit retrieves the built-in portrait shooting scene parameter template T, using the full shooting dataset D and the scene parameter template T as input sources. It builds a five-layer shooting semantic parsing model U based on the core processing unit C. Model U performs temporal analysis and behavior parsing on the motion features of people in dataset D, sorts out the motion state of people at different time points, matches the semantic association logic in template T, extracts the shooting behavior pointing to the motion of people, and finally generates a shooting control strategy V adapted to indoor portrait shooting. Strategy V includes core contents such as subject tracking, lens parameter adjustment, and shooting process control. After receiving the shooting control strategy V, the parameter scheduling unit first breaks down the strategy into four specific control requirements: scene adaptation control, subject tracking control, camera movement trajectory control, and lens parameter control. Then, it performs a logic compliance check. The check follows three fixed rules: using the physical limits of the shooting device (A) to check whether the parameters exceed the rated range; using kinematic constraints (B) to check whether the dynamic operating parameters of the device are compliant; and using scene safety rules (C) to check whether the control commands will cause device collisions or exceed limits. At the same time, it follows the standardized communication protocol (P) to complete the command adaptability check and eliminate all non-compliant control content. After the check passes, the scheduling unit converts the strategy into hierarchical control logic commands (W) that can be recognized by each execution component of the shooting device. After receiving the instruction W, the adaptive execution unit drives the shooting device to complete the structural adjustments of camera positioning, gimbal attitude adjustment, and lens parameter settings according to the instruction requirements. At the same time, it manages the entire process of shooting start, image recording, and shot maintenance according to the preset rhythm. During operation, the execution unit collects the actual effect features E of the shooting image in real time at a fixed sampling frequency F. It compares the features E with the preset shooting target requirements G in the template T in all dimensions one by one, and calculates the deviation value X according to the multi-dimensional vector weighted fusion method: first, it calculates the independent deviations of each dimension of composition, focus, subject tracking, and image brightness to form a multi-dimensional vector, and then integrates them into a single comprehensive deviation scalar through weighted summation to fully reflect the difference between the actual shooting effect and the target requirements. When the overall deviation scalar X exceeds the preset allowable deviation threshold Z, the system automatically triggers an adaptive adjustment process. A tiered priority adjustment algorithm is used to correct the parameters of the control logic command W. The correction magnitude follows the preset parameter Y, prioritizing adjustments to commands related to subject tracking and lens parameters, followed by optimization of scene adaptation parameters. After the command correction is complete, the system waits a fixed duration H for the adjustment effect to take effect, re-acquires the actual image effect features, and recalculates the deviation value X, repeating the iterative correction operation. When the overall deviation scalar X falls back to within the allowable deviation threshold Z, the iterative optimization termination condition is met. The system immediately locks the current control parameters and the operating status of the shooting device, continuously monitoring the shooting status and feature changes at a sampling frequency F to maintain stable portrait shooting effects and complete the closed-loop operation of the entire process.
[0042] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A shooting control system based on pre-made scene templates and structured parameter settings, characterized in that, include: The spatial perception unit establishes a unified global spatial benchmark based on the acquired full-dimensional features from the images. Spatial coordinate calibration and temporal synchronization are performed on various features. The calibrated and synchronized features are then classified and integrated to form a structured full dataset of the captured images. The semantic modeling unit, based on the full dataset of the captured images and combined with pre-made scene parameter templates, constructs a shooting semantic parsing model; performs temporal analysis and behavior parsing on the motion characteristics of the subject being shot, obtains the shooting semantics corresponding to the subject's motion, and generates corresponding shooting control strategies; The parameter scheduling unit performs a full-dimensional breakdown and logic compliance verification of the shooting control strategy; and generates directly executable hierarchical control logic instructions based on the verified strategy according to the shooting device control logic. The adaptive execution unit performs structural adjustments to the shooting device and controls the entire shooting process according to the control logic instructions; it compares the actual effect characteristics of the real-time acquired images with the preset target requirements in all dimensions; it calculates the deviation value between the actual effect characteristics and the target requirements, dynamically corrects instructions according to the degree of deviation value, and iteratively optimizes the actions and related parameters of the shooting device.
2. The system of claim 1, wherein, The specific process for establishing a unified global spatial benchmark is as follows: Obtain the physical spatial layout of the shooting scene, select a fixed point in the scene as the origin of the coordinate system, and preset the direction of the coordinate axis extension and the coordinate range covering the effective shooting area. Selected reference points are spatially located and marked, and a three-dimensional spatial coordinate system for the shooting scene is constructed based on the marked points; The three-dimensional spatial coordinate system is adjusted based on the shooting conditions to match the range of motion of the subject being shot with the operating range of the shooting device.
3. The system of claim 1, wherein, The specific process of performing spatial coordinate calibration and temporal synchronization processing on various features is as follows: Extract the original spatial information during the collection of various features, map the spatial data of various features to the coordinate system of the global spatial reference, correct the spatial position deviations caused by various features, and assign a unified spatial coordinate identifier to various features; Extract the collection time information of various features, align the timelines of different feature data with a unified timeline as the benchmark, fill in the missing time nodes of different feature data, and remove redundant and invalid data on the timeline.
4. The system of claim 1, wherein, The specific process for forming the full dataset of the captured images is as follows: Distinguish between the spatial characteristics of the shooting scene, the motion characteristics of the subject being shot, and the execution state characteristics of the shooting control; Distinguish between each type of core feature and standardize the format of different feature data; Obtain the data storage format and dimensional relationship of various features, and establish mutual retrieval relationships for different feature data.
5. The system of claim 1, wherein, The semantic parsing model for image capture includes a template parsing layer, an architecture determination layer, a data import layer, a model training layer, and a logic optimization layer. The specific construction process is as follows: The input consists of the full dataset of images captured and pre-made scene parameter templates; the output is the image semantic parsing model. The template parsing layer retrieves pre-made scene parameter templates and parses the preset shooting parsing dimensions, feature analysis logic, and shooting semantic association rules; the architecture determination layer determines the overall architecture and core output direction based on the parsing results; the data import layer imports the full shooting dataset and integrates the feature association relationship between the shooting scene and the shooting subject; The model training layer continuously trains parsing capabilities by combining feature association relationships; the logic optimization layer optimizes the feature matching logic.
6. The system of claim 1, wherein, The specific process of performing time-series analysis and behavior analysis on the motion characteristics of the subject being photographed is as follows: Extract motion feature data of the subject being photographed, remove invalid interference information from the data, and form continuous time series data of the subject's motion; The analysis focuses on the direction, amplitude, and rhythm of the subject's movement at different time stages. Based on the semantic association logic of the scene parameter template, the shooting behavior corresponding to each motion state of the subject is analyzed to obtain the core shooting semantics corresponding to the overall motion of the subject.
7. The system of claim 1, wherein, The specific process of conducting full-dimensional decomposition and logical compliance verification is as follows: The shooting control strategy is broken down into specific control requirements for scene adaptation control, subject tracking control, camera movement trajectory control, and lens parameter control. To obtain the actual execution capabilities of the shooting device, and based on the actual spatial conditions of the shooting scene, to check each of the specific control requirements after disassembly, and to eliminate specific control requirements that do not match the equipment capabilities and scene conditions.
8. The system of claim 1, wherein, The specific process for generating directly executable hierarchical control logic instructions is as follows: The hierarchical division and dedicated control logic of each component of the preset shooting device; The verified control requirements for each dimension are allocated according to the execution component level and converted into control information that can be recognized by the corresponding execution component. The main control instructions and auxiliary control instructions are divided according to the importance of the control requirements and the execution order, and then integrated to form a hierarchical control logic instruction.
9. The system of claim 1, wherein, The specific process of structural adjustment of the shooting device and full-process control of the shooting process is as follows: Perform hierarchical parsing of control logic instructions to obtain the adjustment requirements and action ranges corresponding to each execution component; The actuators of the shooting device are driven to perform corresponding adjustment actions, and the adjustment status of each actuator is monitored in real time. Based on the preset shooting rhythm and execution sequence, control the entire process of starting and stopping shooting, switching shot sizes, and adjusting camera movements.
10. The system of claim 1, wherein, The specific process of comparing the real-time acquired image effect features with the preset target requirements in all dimensions is as follows: Real-time capture of the captured footage, and extraction of actual effect features from the real-time captured footage; For each actual effect feature, quantitative extraction and corresponding characterization are performed. Based on the preset shooting target requirements, the corresponding judgment criteria and quantitative indicators for each dimension are obtained. The quantitative data of actual features are compared with the corresponding judgment criteria, and the comparison results and differences are recorded.
11. The system of claim 1, wherein, The specific process for calculating the deviation between the actual effect characteristics and the target requirements is as follows: For comparison dimensions where differences exist, the degree of difference between the actual effect characteristics and the target requirements is quantitatively calculated according to the preset unified evaluation standard. Based on the degree of influence of each dimension on the overall shooting, a weight ratio is set for different comparison dimensions; The weighted calculation of the degree of difference in each dimension is performed, and the weighted difference values of all dimensions are integrated and summarized. A comprehensive deviation value of the overall shooting effect is formed according to a unified calculation rule, and the specific deviation value of each single dimension is extracted and retained.
12. The system of claim 1, wherein, The specific process of dynamically correcting commands and iteratively optimizing the actions and related parameters of the shooting device is as follows: The preset deviation value classification standard determines the deviation level of the overall shooting effect; Develop corresponding instruction correction schemes for different deviation levels, and locate the cause of shooting deviation and the corresponding control dimension based on the single-dimensional deviation value; Adjust the relevant parameters of the corresponding dimension in the control logic command, send the corrected control logic command to the shooting device, and iteratively adjust the actions and parameters of the shooting device.