Camera, camera parameter adjustment method and system
By analyzing the user's shooting intentions and scene perception, and utilizing a multi-parameter collaborative decision-making intelligent network model for photography, the system achieves automated multi-parameter collaborative optimization of camera parameters. This solves the problems of complex operation and poor results in existing technologies, and improves the shooting quality for non-professional users.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SMARTSENS TECH (SHANGHAI) CO LTD
- Filing Date
- 2026-06-26
- Publication Date
- 2026-07-31
AI Technical Summary
Existing camera shooting parameter adjustment solutions have high operational barriers and low efficiency. They cannot interpret the user's shooting intentions at the effect level, lack scene awareness, cannot achieve multi-parameter collaborative decision-making, and lack scene-based parameter constraint rules, resulting in poor shooting effects.
By parsing the user's voice or text shooting intention command, a shooting intention feature vector is generated. Combined with the camera preview image perception feature vector, a set of camera adjustment parameters is generated using a photography multi-parameter collaborative decision-making intelligent network model to achieve multi-parameter collaborative optimization.
Users can achieve professional-level parameter adjustments without any professional knowledge, adapt to the current scene, improve shooting results, unleash the performance of camera hardware, and eliminate the professional barrier to photography.
Smart Images

Figure CN122496712A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of camera parameter adjustment technology, and in particular to a camera, camera parameter adjustment method and system. Background Technology
[0002] Current camera shooting parameter adjustment methods mainly rely on physical buttons, dials, and touch screen menus to manually set multiple parameters such as shutter speed, aperture, ISO, and white balance, which is difficult to operate and inefficient. Existing camera voice control solutions (such as CN115297237A, CN110381297A, and CN114500890A) focus primarily on optimizing voice recognition accuracy and wake-up mechanisms, supporting only basic commands such as taking photos, recording videos, and adjusting single parameters. These solutions suffer from four unavoidable core technical flaws, creating a significant technological gap: 1. Unable to interpret the intended effect of shooting, and the demand and parameters are completely disconnected: The existing solution can only recognize standardized parameter operation commands and cannot interpret the user's spoken effect-level demand (such as "shoot Tyndall effect", "darken the environment to highlight the subject", "make the skin more transparent"). Ordinary users cannot convert shooting effect demands into professional parameter commands. Voice control can only serve users with professional photography knowledge and cannot achieve universal application. 2. Lack of scene awareness and unadaptable parameter adjustments, easily causing irreversible image quality loss: Existing solutions can only perform fixed adjustments to a single parameter according to instructions, without considering the environment, subject, and imaging constraints of the shooting scene. For example, if a user requests to "reduce brightness," the existing solution will directly reduce the aperture, which will directly cause the background blur to disappear in portrait scenes; if a user requests to "reduce noise," the existing solution will directly reduce the ISO and increase the shutter speed, which will cause subject blur in motion scenes, making it completely impractical. 3. Lack of multi-parameter collaborative decision-making capability: The core parameters of photography (aperture, shutter speed, ISO, etc.) have a strong coupled linkage logic. Adjusting a single parameter will inevitably lead to neglecting other aspects. Existing solutions lack the ability to optimize multi-parameter collaboratively, and cannot meet user needs while simultaneously ensuring image quality, effect integrity, and optimal hardware performance. 4. Lack of scenario-based parameter constraints leads to adjustment logic that violates photographic principles: The priority and constraint boundaries of parameter adjustments are completely different in different scenarios. Existing solutions lack scenario-based constraint rules, making it impossible to adapt to the needs of shooting in all scenarios and making it easy for parameter adjustments to violate the basic logic of photography. Summary of the Invention
[0003] To address the aforementioned technical issues, the purpose of this application is to provide a camera, a camera parameter adjustment method, and a system that can directly transform a user's shooting needs into a multi-parameter collaborative optimization scheme adapted to the current scene, thereby eliminating the professional barrier to photography.
[0004] To achieve the above objectives: Firstly, this application provides a method for adjusting camera parameters, including: In response to the shooting intention command, the shooting intention command is parsed to generate a shooting intention feature vector; Acquire a camera preview image, identify and quantize the camera preview image, and generate a scene-aware feature vector; Collect the camera initial parameter feature vector and the camera hardware capability boundary feature vector, combine them with the shooting intention feature vector and the scene perception feature vector, and generate a set of camera adjustment parameters based on the photography multi-parameter collaborative decision-making intelligent network model. Configure the camera parameters according to the camera adjustment parameter set, and update the camera preview screen synchronously.
[0005] Furthermore, prior to the step of responding to the shooting intention command, the following steps are included: Obtain the user's voice input shooting intention command, and convert the voice shooting intention command into a text shooting intention command.
[0006] Furthermore, the step of parsing the shooting intention command and generating a shooting intention feature vector includes: The pre-trained natural language understanding model parses the shooting intention command and extracts three types of structured information: intention label, adjustment range, and constraint conditions. The structured information is encoded to generate a shooting intent feature vector that includes intent label encoding, adjustment amplitude value, and constraint condition encoding.
[0007] Furthermore, the step of parsing the shooting intention command and extracting three types of structured information—intention label, adjustment range, and constraint conditions—based on the pre-trained natural language understanding model includes: A four-level shooting intent classification system is constructed, with the levels subdivided from top to bottom: first-level classification, second-level classification, third-level classification, and fourth-level classification. The shooting intent command is matched hierarchically from top to bottom against the four-level shooting intent classification system to determine the classification category with the highest matching degree, so as to generate the intent tag.
[0008] Further, the step of generating the intent label is followed by: The intent confidence value of the intent label is obtained based on the principle of classification probability normalization; When the intent confidence value is lower than the preset intent confidence threshold, scene information is added to fuse the shooting intent command, and the intent tag is regenerated based on the updated shooting intent command.
[0009] Furthermore, the steps of acquiring the camera preview image, identifying and quantizing the camera preview image, and generating a scene-aware feature vector include: The camera preview image is identified and quantified based on a pre-trained computer vision model, and scene information is extracted, including scene labels, environmental imaging parameters, and the subject being photographed. The scene information is encoded to generate a scene perception feature vector containing scene label encoding, environmental imaging parameter encoding, and subject encoding.
[0010] Furthermore, the step of identifying and quantifying the camera preview image based on the pre-trained computer vision model, and extracting scene information, including scene labels, environmental imaging parameters, and the subject being photographed, includes: Construct a first preset number of scene types, match the camera preview with each scene type, identify the scene type with the highest matching degree, and generate the scene tag; The pixel data in the camera preview image is statistically analyzed and quantified to generate the environmental imaging parameter information; A second preset number of subject feature types are constructed, and the camera preview image is matched with each subject feature type one by one to determine the corresponding feature value in order to generate the subject to be photographed.
[0011] Furthermore, the step of constructing a first preset number of scene types, matching the camera preview with each scene type, and identifying the scene type with the highest matching degree to generate the scene label further includes: The scene confidence value for each scene type is obtained based on the principle of classification probability normalization. When the scene confidence value of any scene type is not lower than the preset scene confidence threshold, the scene type is confirmed as the scene type with the highest matching degree, and the scene label is generated.
[0012] Furthermore, the photographic multi-parameter collaborative decision-making intelligent network model includes an input layer, a fusion layer, a feature encoding layer, and an output layer; The step of collecting the camera initial parameter feature vector and the camera hardware capability boundary feature vector, and combining them with the shooting intention feature vector and the scene perception feature vector to generate a camera adjustment parameter set based on a multi-parameter collaborative decision-making intelligent network model for photography includes: The input layer inputs the camera initial parameter feature vector, the camera hardware capability boundary feature vector, the shooting intention feature vector, and the scene perception feature vector; The fusion layer fuses the shooting intent feature vector and the scene awareness feature vector, enabling the camera adjustment parameter set to associate the shooting intent command with the camera preview screen; The feature encoding layer is based on a transform encoder structure and learns the linkage logic in the initial parameters of the camera to adjust the camera adjustment parameter set. The output layer constrains the camera hardware capability boundary feature vector and outputs the camera adjustment parameter set.
[0013] Furthermore, the step of generating the camera adjustment parameter set based on the intelligent network model for multi-parameter collaborative decision-making in photography includes the following prior steps: Construct a pre-trained dataset to train the photographic multi-parameter collaborative decision-making intelligent network model; The multi-task weighted evaluation rule is adopted to constrain and optimize the photography multi-parameter collaborative decision-making intelligent network model from four dimensions: parameter accuracy, intent matching degree, scene adaptability, and rule compliance.
[0014] Furthermore, the step of generating the camera adjustment parameter set based on the intelligent network model for multi-parameter collaborative decision-making in photography also includes: Each of the aforementioned scene types is configured with corresponding camera parameter priority sorting and camera parameter hard constraint rules. The photography multi-parameter collaborative decision-making intelligent network model generates and optimizes the camera adjustment parameter set based on the camera parameter priority sorting and the camera parameter hard constraint rules.
[0015] Further, after the step of configuring the camera parameters according to the camera adjustment parameter set and synchronously updating the camera preview image, the following steps are included: Obtain user feedback information, and train and iteratively optimize the photography multi-parameter collaborative decision-making intelligent network model based on the feedback information.
[0016] Secondly, this application provides a camera parameter adjustment system, including: The shooting intention input module is used to obtain the user's voice input shooting intention command and convert the voice shooting intention command into a text shooting intention command; The intent parsing module is used to parse the shooting intent command and generate a shooting intent feature vector; The scene perception module is used to acquire camera preview images, identify and quantize the camera preview images, and generate scene perception feature vectors. The multi-parameter collaborative decision-making intelligent module is used to collect the camera initial parameter feature vector and the camera hardware capability boundary feature vector, and combine the shooting intention feature vector and the scene perception feature vector to generate a set of camera adjustment parameters based on the photography multi-parameter collaborative decision-making intelligent network model. The parameter execution module is used to configure the camera parameters according to the camera adjustment parameter set and to update the camera preview screen synchronously. The self-learning optimization module is used to acquire user feedback information and train and iteratively optimize the photography multi-parameter collaborative decision-making intelligent network model based on the feedback information. Thirdly, this application provides a camera, including: a processor and a memory storing a computer program, wherein when the processor runs the computer program, it implements the steps of the aforementioned camera parameter adjustment method.
[0017] The camera, camera parameter adjustment method, and system of this application analyze the user's shooting intention command, perceive the current shooting scene, and directly transform the user's shooting needs into a multi-parameter collaborative optimization scheme adapted to the current scene based on a photography multi-parameter collaborative decision-making intelligent network model, thus eliminating the professional threshold of photography. Attached Figure Description
[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without any creative effort.
[0019] Figure 1 This is a schematic flowchart of the camera parameter adjustment method provided in the first embodiment of this application.
[0020] The realization of the objectives, functional features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0021] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0022] It should be understood that although the terms first, second, third, etc., may be used in this document to describe various information, elements, units, or modules, these information, elements, units, or modules should not be limited to these terms. These terms are only used to distinguish information, elements, units, or modules of the same type from one another.
[0023] It should be noted that step designations such as S11 and S12 are used in this document for the purpose of more clearly and concisely describing the corresponding content, and do not constitute a substantial limitation on the order. In specific implementation, those skilled in the art may execute S12 first and then S11, etc., but these should all be within the protection scope of this application.
[0024] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0025] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0026] First Embodiment See Figure 1 The first embodiment of this application (hereinafter referred to as "this embodiment") provides a camera parameter adjustment method, including: S10: In response to the shooting intention command, parse the shooting intention command and generate a shooting intention feature vector.
[0027] For example, user-issued shooting intention commands, such as "I want to increase the brightness" or "I want to shoot Tyndall effect", are parsed end-to-end based on a pre-trained and finely tuned Natural Language Understanding (NLU) model specific to the photography domain. The structured shooting intention feature vector is then extracted and used as input to the intelligent network model for multi-parameter collaborative decision-making in photography.
[0028] S20: Acquire the camera preview image, identify and quantize the camera preview image, and generate a scene-aware feature vector.
[0029] For example, the preview image output by the camera image sensor is acquired in real time. Then, a computer vision (CV) model optimized for the photography scene is used to perform pixel-level recognition and quantization calculation on the preview image, and output a scene-aware feature vector that is strongly bound to the camera adjustment parameter decision.
[0030] S30: Collect the camera initial parameter feature vector and the camera hardware capability boundary feature vector, combine them with the shooting intention feature vector and the scene perception feature vector, and generate a camera adjustment parameter set based on the photography multi-parameter collaborative decision-making intelligent network model.
[0031] For example, the shooting intent feature vector, scene awareness feature vector, current initial parameter set of the camera, and camera hardware capability boundary features are input into a pre-trained Photography Multi-Parameter Collaborative Decision Network (PMPCD-Net) model, which outputs a set of camera adjustment parameters that are adapted to the current scene and match the user's shooting intent.
[0032] S40: Configure the camera parameters according to the camera adjustment parameter set, and update the camera preview screen synchronously.
[0033] For example, the camera adjustment parameter set is aligned with the camera image sensor's frame synchronization signal (VSYNC, Vertical Synchronization) to achieve simultaneous frame-by-frame adjustment of hardware imaging parameters and ISP (Image Signal Processor) image processing parameters. All parameters take effect simultaneously in the same frame, avoiding screen flicker and updating the camera's preview in real time. After the camera parameters are adjusted, the changes in core parameters are displayed on the camera screen.
[0034] In this embodiment, the camera parameter adjustment method parses the user's shooting intention command, perceives the current shooting scene, and directly transforms the user's shooting needs into a multi-parameter collaborative optimization scheme adapted to the current scene based on a multi-parameter collaborative decision-making intelligent network model. Users do not need any professional photography knowledge; they can obtain a professional-grade parameter scheme adapted to the scene through simple shooting intention commands, significantly improving the success rate of photos taken by non-professional users, fully unleashing the camera hardware performance, and eliminating the professional barrier to photography.
[0035] In one embodiment, S10: prior to the step of responding to the shooting intention command, includes: Obtain the user's voice input shooting intention command, and convert the voice shooting intention command into a text shooting intention command.
[0036] For example, spoken shooting instructions from the user are converted into corresponding text-based shooting intent statements. In this embodiment, the user does not need any professional photography knowledge; they can obtain a professional-grade parameter scheme suitable for the scene simply by expressing their desired effect. In another embodiment, the user directly inputs shooting intent instructions to the camera in text; this application does not limit this approach.
[0037] In one embodiment, S10: The step of parsing the shooting intention command and generating a shooting intention feature vector includes: S11: Based on a pre-trained natural language understanding model, the shooting intention instruction is parsed, and three types of structured information are extracted: intention label, adjustment range, and constraint conditions.
[0038] For example, a natural language understanding model pre-trained in the field of photography is used to extract three core pieces of structured information. The intent label is the user's actual intent requirement in the shooting intent command, the adjustment range is the fuzziness level description in the shooting intent command (such as "a little", "a lot", "full", etc.), and the constraint condition is the restrictive requirement specified by the user in the shooting intent command (such as "do not destroy the bokeh" "keep the image clear" etc.).
[0039] S12: Encode the structured information to generate a shooting intent feature vector containing intent label encoding, adjustment amplitude value, and constraint condition encoding.
[0040] For example, the intent label is matched with multiple photography intent categories in the natural language understanding model, and the intent label code is generated based on the intent category with the highest matching degree. The adjustment range is uniformly quantized into 7 standardized values from -3 to +3, where negative values represent reduction or weakening, and positive values represent increase or enhancement. The quantized values are then encoded into floating-point and / or integer values according to the symmetric quantization rule to generate the adjustment range value. The constraints are converted into hard constraint rules for parameter adjustment, corresponding to various constraint rules already stored in the natural language understanding model. This encoding segment is a binary and / or one-hot encoding area of a preset length. If no user-specified constraint exists, all constraint bits are uniformly set to "invalid bit (0)". If a constraint exists, the corresponding constraint bit is set to 1, and the rest are 0. Finally, a 64-dimensional standardized shooting intent feature vector is output, which includes the unique encoding of the intent label, the adjustment range value, and the constraint condition encoding. Optionally, this application does not limit the dimension of the shooting intent feature vector.
[0041] In this embodiment, the camera parameter adjustment method completely solves the core pain point of the disconnect between requirements and parameters: through photography-specific intent parsing technology, it achieves accurate conversion of spoken effect requirements into structured shooting intentions, with a spoken requirement matching accuracy rate of no less than 93% and a composite requirement matching accuracy rate of no less than 86%, which is completely different from the shortcomings of existing solutions that can only recognize fixed parameter instructions.
[0042] In one embodiment, S11: The step of parsing the shooting intention instruction based on a pre-trained natural language understanding model and extracting three types of structured information—intention label, adjustment range, and constraint conditions—includes: S13: Construct a four-level shooting intent classification system, with the levels gradually subdivided from top to bottom, namely, first-level classification, second-level classification, third-level classification and fourth-level classification.
[0043] For example, a predefined four-level intent classification system for photography is defined. The first-level classification includes four major categories, which are further broken down into 16 second-level categories, 48 third-level category sub-items, and 128 fourth-level category effect tags, achieving full coverage of the needs for conversational shooting in all scenarios.
[0044] Specifically, the first-level categories (4 categories): exposure control, image quality optimization, portrait effects, and scene effects, dividing the field into four core photography needs domains, covering all major categories of shooting intentions; the second-level categories (16 categories): each first-level category is equally divided into 4 second-level subcategories (4×4=16), completing the medium-granularity demand breakdown; the third-level categories (48 categories): each second-level category is equally divided into 3 third-level sub-items (16×3=48), corresponding to specific shooting actions / demands that can be implemented; the fourth-level categories (128 categories): based on the third-level categories, further subdivided into effects, styles, and local limitation tags.
[0045] Optionally, this application does not limit the number of categories at each level. Users can design the number of categories at each level according to their photographic intent, as long as the hierarchical relationship remains unchanged (level 1 → level 2 → level 3 → level 4) and the tag system is mutually exclusive and fully covered.
[0046] S14: Match the shooting intent command with the four-level shooting intent classification system from top to bottom, determine the classification category with the highest matching degree, and generate the intent tag.
[0047] For example, the shooting intent command is matched against a four-level shooting intent classification system, and the category with the highest matching degree is obtained at each level, which is then combined to generate an intent tag. The encoding value of each category in each level of the shooting intent classification system is defined, and the encoding value of each category in the intent tag is obtained, which is then combined to generate an intent tag code. For example, a fixed-length segmented encoding scheme is adopted, with each level of the four-level classification uniformly occupying 2 decimal digits, and the total length of the complete tag code is fixed at 8 bits. The coding value range and placeholder rules for each level are as follows: (1) First-level major category coding: value range 01~99, covering all top-level shooting scene major categories; (2) Second-level intermediate category coding: value range 01~99, corresponding to the subdivided shooting scenes under the first-level major category, filled with 00 when there is no second-level subdivision; (3) Third-level minor category coding: value range 01~99, corresponding to the subdivided shooting style and / or object under the second-level intermediate category, filled with 00 when there is no third-level subdivision; (4) Fourth-level detailed category coding: value range 01~99, corresponding to the refined shooting needs under the third-level minor category, only the categories split to the third level are uniformly filled with placeholder code 00. Therefore, the four-level shooting intent classification system is layered and the category coding is assigned. After matching the optimal category at each level, the corresponding segment code is filled in and spliced into the complete code, i.e., the intent label code.
[0048] Understandably, there are cases where third-level categories are not split into fourth-level categories, and there are also cases where shooting intent commands (such as "too bright" or "reduce noise") are general basic intents, and the third-level categories can fully describe the requirements, so there is no need to attach fourth-level categories. Therefore, when generating intent tag codes, if there is a matching category, the corresponding category code value will be filled in; if there is no fourth-level category, the field will be filled with a preset default value (such as an invalid tag code or a placeholder).
[0049] In this embodiment, a four-level shooting intent classification system is constructed, with the levels progressively subdivided from top to bottom to reduce the difficulty of recognizing shooting intent commands. The system first completes the classification of major categories (coarse-grained) and then further subdivides them into third-level and fourth-level categories (fine-grained). This avoids the natural language understanding model from directly searching the massive number of fine tags globally, thereby improving the inference speed and accuracy of the model.
[0050] In one embodiment, different secondary categories can be pre-bound with parameter adjustment information, including the general direction of parameters, basic weights, and pre-constraints. For example, the "exposure control" category in the primary category can be subdivided into secondary categories such as overall brightness adjustment, local brightness adjustment, and dynamic range adjustment. Different secondary categories directly connect to different parameter groups, reducing the computational load of subsequent decision networks. Specifically, the secondary category, as a functional partitioning layer, is responsible for dividing independent imaging adjustment functional domains (such as exposure, color, blur, sharpening, etc.). The underlying parameter groups of different adjustment functions are completely isolated. Therefore, the general direction of parameters, basic weights, and pre-constraints in the parameter adjustment information are bound to each category of the secondary category, realizing a one-to-one mapping between functions and parameter pools, pre-distributing parameter logic, and reducing the computational load of subsequent networks.
[0051] In one embodiment, the third-level and fourth-level classifications serve as local scene subdivision layers, responsible for subdividing and fine-tuning local scene features. Each category is pre-bound with adjustment amplitude and subdivision constraints from parameter adjustment information. During actual matching, the parameter adjustment information pre-bound to each category at levels two, three, and four is fused based on a layered overlay binding principle. For example, categories at levels three and four inherit the basic parameter adjustment information of categories at level two and overlay subdivision correction parameters, without overwriting the core parameter adjustment information pre-bound to categories at level two. Categories at level three are pre-bound with subdivision style correction weights and local parameter offsets; categories at level four are pre-bound with refined scene thresholds and fine-tuning amplitude offsets. Understandably, the underlying basic parameter pool is still uniquely connected to categories at level two, and categories at levels three and / or four only undergo fine-tuning corrections without changing the basic parameter group.
[0052] In one embodiment, S14: After the step of generating the intent label, the following is included: The intent confidence value of the intent label is obtained based on the principle of classification probability normalization.
[0053] For example, when the category with the highest matching degree at each level is output, the confidence value of each category is obtained and combined to generate the intent confidence value of the intent tag. Specifically, the intent confidence is calculated based on the normalization of classification probability and the textual semantic features of the shooting intent command, corresponding to four levels of shooting intent classification categories, and is used to determine the reliability of the intent recognition result. The intent confidence value adopts a scheme of independent calculation at each level and outputting probabilities layer by layer. The specific calculation logic is as follows: Calculation dimension: The confidence is calculated independently at the first to fourth level classification levels. Each level only performs semantic matching scoring within its own category set, without interference. Finally, the probability value corresponding to the category with the highest matching degree at each level is taken as the confidence value of that level. Calculation method: Based on the pre-trained semantic matching model, the semantic similarity between the user's shooting intention command and the description text of each level category is calculated. After normalization by the softmax function, it is mapped to the probability value in the range of 0 to 1, which is the confidence value of the corresponding category. Threshold adaptation: 0.85 is the verification threshold for sub-categories (level 3 and level 4). Level 1 and level 2 categories can be set with lower thresholds or exempted from verification because of their coarse semantic granularity and high distinguishability.
[0054] When the intent confidence value is lower than the preset intent confidence threshold, scene information is added to fuse the shooting intent command, and the intent tag is regenerated based on the updated shooting intent command.
[0055] For example, when the category confidence value of the third-level and / or fourth-level classification is lower than the preset intent confidence threshold of 0.85, scene information is added to trigger secondary confirmation. Specifically, when the confidence value of the fine-grained third-level and / or fourth-level classification is insufficient, secondary confirmation is triggered first. If the confidence value is still insufficient, it can fall back to the parameter group of the second-level classification to execute general adjustment logic, thereby improving the robustness of the natural language understanding model. Optionally, this application does not limit the size of the preset intent confidence threshold; users can set it themselves according to the accuracy of the matching. Preferably, the preset intent confidence threshold is 0.85.
[0056] For example, there are two ways to perform secondary confirmation: (1) The semantic verification and completion module on the end side does not require user participation throughout the process. The specific execution process is as follows: 1. When the confidence of the category of the third-level and / or fourth-level classification is lower than 0.85, a second matching is triggered: supplement the image features of the current view, that is, the scene information includes the scene category, the subject type, the exposure level, etc., and perform multimodal fusion with the original text command; 2. Reorder the matching in the category candidate set of the third-level and / or fourth-level classification, and output the new optimal category and the corresponding confidence; 3. If the confidence after the second calculation is not lower than 0.85, then the sub-category and the corresponding sub-sub-parameters are adopted; if it is still lower than the threshold, then it automatically falls back to the general parameter group corresponding to the category of the second-level classification to perform adjustment. After the number of fallbacks N, the confidence still cannot be lower than 0.85, then the scheme (2) is executed.
[0057] (2) Initiate an interactive confirmation with the current shooting user. The specific execution process is as follows: 1. When the confidence of the sub-category is insufficient, 2-3 high-probability third-level and / or fourth-level category candidate intent options will pop up on the camera interaction interface, with a brief description; 2. After the user manually selects the target category, the corresponding sub-parameters will be loaded directly; if the user does not operate within the time limit or selects "default adjustment", the system will automatically revert to the general parameter logic of the second-level category.
[0058] In another embodiment, if the shooting intention command does not contain structured information on adjustment range and / or constraints, a shooting intention feature vector is generated based on the parameter adjustment information pre-bound to the most granular category matched in the intention tag. If the shooting intention command contains structured information on adjustment range and / or constraints, a shooting intention feature vector is generated by combining the parameter adjustment information pre-bound to the most granular category matched in the intention tag. Specifically, two mutually exclusive and fully covered generation branches can be set, corresponding to two types of user input scenarios: (1) Scenarios without explicit structured parameters (users only specify style and / or scenario, such as "I want Tyndall effect") When the shooting intention command does not extract clear adjustment range, constraints and other structured parameters, the full set of parameter adjustment information pre-bound to the most granular level category is directly used to generate the shooting intention feature vector.
[0059] (2) Scenarios with explicit structured parameters (users provide quantitative descriptions, such as "brighten a little, don't overexpose") When a clear adjustment range and constraint can be extracted from the shooting intention command, the shooting intention feature vector is generated by superimposing the user-specified range correction and constraint conditions on the basic parameters bound to the most refined category, and then merging them.
[0060] In one embodiment, S20: The step of acquiring a camera preview image, identifying and quantizing the camera preview image, and generating a scene-aware feature vector includes: S21: Based on a pre-trained computer vision model, identify and quantify the camera preview image, and extract scene information, including scene labels, environmental imaging parameters, and the subject being photographed.
[0061] For example, a computer vision model is pre-trained and fine-tuned specifically for photographic scenarios. A lightweight dual-branch model for instance segmentation and scene classification is employed, with the camera preview YUV image as input. Three dimensions of scene information are extracted, with a single-frame inference time of no more than 20ms, without affecting the preview frame rate, ensuring real-time scene recognition. The computer vision model is deployed online or locally / offline on the camera's NPU (Neural Processing Unit) / ISP processor.
[0062] For example, the scene label is the real-time scene type in the camera preview, and scene label information is extracted from the camera preview based on a scene classification model. Environmental imaging parameters are pixel-level parameter data in the camera preview, including ambient EV (Exposure Value), highlight / shadow ratio, dynamic range, contrast, absolute color temperature, light source type, and noise level. The subject is the real-time subject information in the camera preview, and subject information is extracted from the camera image based on a lightweight instance segmentation model, including subject type (face, body, object, etc.), subject position and proportion in the image, number of faces and 68 key points, subject movement speed, focus status, and subject lighting characteristics.
[0063] S22: Encode the scene information to generate a scene perception feature vector containing scene label encoding, environmental imaging parameter encoding, and subject encoding.
[0064] For example, the scene label is matched with multiple scene types already constructed in the scene classification model to obtain the scene type with the highest matching degree, and the encoding value of this scene type is confirmed to generate a 36-dimensional scene label code. After normalizing the environmental imaging parameters to 0 to 1, a 30-dimensional environmental imaging parameter code is generated. The subject being photographed is matched with predefined subject features, and the corresponding feature values are confirmed to generate a 30-dimensional subject code. Finally, a 96-dimensional standardized scene perception feature vector is output, including the scene label one-hot encoding, the environmental imaging parameter normalization matrix, and the subject information encoding. Optionally, this application does not limit the dimension of the scene perception feature vector.
[0065] In one embodiment, S21: The step of identifying and quantifying the camera preview image based on a pre-trained computer vision model, and extracting scene information, wherein the scene information includes scene labels, environmental imaging parameters, and the subject being photographed, includes: S23: Construct a first preset number of scene types, match the camera preview with each scene type, and confirm the scene type with the highest matching degree to generate the scene label.
[0066] For example, a classification system of 10 primary scenes and 36 secondary sub-scenes is constructed. Primary scenes include portrait, landscape, night scene, backlight, indoor, sports, macro, pet, starry sky, and urban architecture. Secondary sub-scenes are the subdivided shooting scenes under the corresponding primary scenes. When matching the camera preview with scene types, the primary scenes are matched first to obtain the primary scene type with the highest matching degree, and then the subdivided secondary sub-scenes under this primary scene type are matched. Optionally, this application does not limit the first preset number; the user sets the first preset number according to scene coverage requirements. In this embodiment, the first preset number is 36. The two-level scene classification system is designed to speed up the matching process.
[0067] S24: Statistically analyze and quantify the pixel data in the camera preview image to generate the environmental imaging parameter information.
[0068] For example, pixel-level statistics are performed on the YUV channels of the camera preview image to output a standardized environmental imaging parameter matrix.
[0069] S25: Construct a second preset number of subject feature types, match the camera preview image with each subject feature type one by one, and determine the corresponding feature value to generate the subject to be photographed.
[0070] For example, the subject information can be accurately extracted based on the lightweight instance segmentation and face detection model. Optionally, this application does not limit the number of the second preset. The user can set the corresponding number of the second preset according to the shooting requirements. In this embodiment, the number of the second preset is 6. Other dimensions such as subject position coordinates, number of subjects, edge contour features and other subject feature types can also be set. The detection results can be assigned during encoding.
[0071] The above examples illustrate the process of generating scene-aware feature vectors, which includes scene label encoding, environmental imaging parameter encoding, and subject encoding: Encoding Segment 1: 36-dimensional secondary scene One-Hot encoding The scene classification system has 36 secondary scenes. The current recognition is an outdoor backlit portrait (assuming it corresponds to the 8th secondary scene). The encoding rule is: only the 8th dimension = 1, and the remaining 35 = 0.
[0072] Encoding Segment 2: 30-Dimensional Environmental Imaging Parameters After extracting and normalizing the statistical values of the image, the core dimensions are selected (mapped from the measured data of the example): Ambient EV value 14.5, normalized to 0.82; Highlight ratio 31%, normalized to 0.31; Shadow ratio 12%, normalized to 0.12; Dynamic range 13EV, normalized to 0.91; Absolute color temperature 5600K, normalized to 0.75; Light source type (backlight, natural light), encoded to 0.95; Image noise level, encoded to 0.10; The remaining dimensions are image contrast, average brightness, color distribution, etc., filled with decimals between 0 and 1 according to the measured values.
[0073] Encoding segment 3: Encoding of subject information from 30-dimensional imaging Current subject: Single person in the center of the image, face occupies 24% of the frame, stationary, face locked in focus. Encoding rules: Subject type (portrait), assigned 1.0; subject occupies 24% of the frame, assigned 0.24; subject movement speed (stationary), assigned 0.0; focus status (precise focus), assigned 0.98; completeness of 68 facial key points, assigned 0.97; subject lighting (backlit face), assigned 0.88; other dimensions such as subject position coordinates, number of subjects, edge contour features, etc., are assigned according to the detection results.
[0074] Based on this, a 96-dimensional scene-aware feature vector can be obtained: [0,0,1 (8th position),0,...0] (first 36-dimensional scene label encoding) + [0.82,0.31,0.12,0.91,0.75,...] (middle 30-dimensional environmental imaging parameter encoding) + [1.0,0.24,0.0,0.98,0.97,...] (last 30-dimensional subject encoding).
[0075] In one embodiment, S23: The step of constructing a first preset number of scene types, matching the camera preview image with each scene type, and identifying the scene type with the highest matching degree to generate the scene label further includes: The scene confidence value for each scene type is obtained based on the principle of classification probability normalization.
[0076] For example, the confidence score is normalized using Softmax: ;z j : The original output score of the model for the i-th scene; n: The total number of scenes to be matched (n=10 for first-level scenes, n=10 for second-level scenes, and n=10 for sub-scenes under the corresponding major category); Conf i : The final confidence level of the i-th scenario (0~1, equivalent to 0%~100%).
[0077] When the scene confidence value of any scene type is not lower than the preset scene confidence threshold, the scene type is confirmed as the scene type with the highest matching degree, and the scene label is generated.
[0078] For example, in the first-level scene matching stage: the scene with the highest confidence among 10 first-level scenes is selected, with no mandatory threshold (mainly used for rapid initial screening before entering the second-level matching pool). In the second-level scene matching stage: the scene confidence value of all scene types in the second level is calculated sequentially. When the scene confidence value of a new scene type is not lower than a preset scene confidence threshold, the scene type is switched to the new scene type, and the scene label and corresponding constraint rules are updated; when the scene confidence value of a new scene type is lower than the preset scene confidence threshold, the scene type of the previous frame is maintained unchanged, and the scene type switch is refused. The scene confidence value is calculated based on the visual features of the viewfinder and corresponds to the shooting scene classification category, used to determine the reliability of the scene recognition result. Optionally, this application does not limit the size of the preset scene confidence threshold; users can set it themselves according to the accuracy of scene type matching. Preferably, the preset scene confidence threshold is 90%.
[0079] In this embodiment, the camera preview image is matched one by one with the above scene classification system to obtain the second-level sub-scene with the highest matching degree. The scene classification matching accuracy based on the scene classification model is not less than 96%. The scene label is updated only when the confidence value of the new scene is not lower than 90% of the preset scene confidence threshold to avoid frequent switching caused by screen flicker. At the same time, the scene recognition results are further corrected by combining the data from the camera gyroscope, focus distance, and ambient light sensor to improve the recognition stability.
[0080] In another embodiment, during the secondary confirmation process when the intent confidence value is lower than a preset intent confidence threshold, the scene confidence value and the corresponding scene label can be used as auxiliary features to participate in the multimodal recalculation of the intent confidence value, so as to improve the accuracy of fine-grained intent recognition.
[0081] In one embodiment, the photographic multi-parameter collaborative decision-making intelligent network model includes an input layer, a fusion layer, a feature encoding layer, and an output layer.
[0082] S30: The step of collecting the camera initial parameter feature vector and the camera hardware capability boundary feature vector, and combining them with the shooting intention feature vector and the scene perception feature vector to generate a camera adjustment parameter set based on a multi-parameter collaborative decision-making intelligent network model includes: The input layer takes into account the camera initial parameter feature vector, the camera hardware capability boundary feature vector, the shooting intention feature vector, and the scene perception feature vector.
[0083] For example, the input layer is a fixed 192-dimensional feature vector, including a 64-dimensional shooting intention feature vector, a 96-dimensional scene awareness feature vector, a 20-dimensional camera initial parameter feature vector, and a 12-dimensional camera hardware capability boundary feature vector (such as lens aperture range, shutter speed limit, native ISO range, image stabilization capability, etc.). Optionally, this application does not limit the dimensions of the above feature vectors.
[0084] The fusion layer fuses the shooting intent feature vector and the scene awareness feature vector, so that the camera adjustment parameter set is associated with the shooting intent command and the camera preview screen.
[0085] For example, the fusion layer includes two layers of intent and scene cross-attention modules. Through a multi-head cross-attention mechanism, it learns the relationship between shooting intent and scene features, thus solving the core problem of the disconnect between needs and scenes.
[0086] The feature encoding layer is based on a transform encoder structure and learns the linkage logic in the initial parameters of the camera to adjust the camera adjustment parameter set.
[0087] For example, the feature encoding layer includes a 3-layer Transformer Encoder structure, which learns the coupling logic between photographic parameters through a self-attention mechanism to resolve the conflict problem of multi-parameter adjustment.
[0088] The output layer constrains the camera hardware capability boundary feature vector and outputs the camera adjustment parameter set.
[0089] For example, the output layer is a dual-branch parallel output, covering the entire camera parameter chain, with all output parameters clamped to the legal range supported by the camera hardware. One layer is the hardware imaging parameter branch: outputting 7 core hardware parameters including shutter speed, aperture value, ISO sensitivity, white balance color temperature, focus position, focus mode, and image stabilization mode; the other layer is the ISP image processing parameter branch: outputting 9 core ISP parameters including exposure compensation, noise reduction intensity, sharpening intensity, contrast, saturation, tonal curve parameters, beauty parameters, and effects algorithm on / off and intensity.
[0090] In this embodiment, the core architecture of the PMPCD-Net model is an end-to-end neural network customized for dedicated photography parameter linkage optimization. It is based on the Transformer-Encoder architecture and adds an intent-scene cross-attention fusion module to achieve deep coupling between shooting requirements and scene constraints. The size of the full model after INT8 quantization is no more than 80MB. It can be deployed offline locally, and the single inference time is no more than 20ms, which is compatible with the camera's low-power processor.
[0091] In one embodiment, S30: Before the step of generating a camera adjustment parameter set based on a photographic multi-parameter collaborative decision-making intelligent network model, the following steps are included: A pre-trained dataset is constructed to train the photographic multi-parameter collaborative decision-making intelligent network model.
[0092] For example, three photography-specific pre-training datasets are constructed, with a total size of no less than 10 million sets, including more than 2 million sets of scene-based parameter adjustment samples from professional photographers, more than 6 million sets of "shooting intention-scene-optimal parameters" mapping samples, and more than 2 million sets of scene-parameter annotation samples from public photography datasets.
[0093] The multi-task weighted evaluation rule is adopted to constrain and optimize the photography multi-parameter collaborative decision-making intelligent network model from four dimensions: parameter accuracy, intent matching degree, scene adaptability, and rule compliance.
[0094] For example, the multi-task weighted evaluation rule is based on the multi-task weighted loss function: The total loss is calculated as αLparam + βLintent + γLscene + δLconstraint, where α = 0.4, β = 0.3, γ = 0.2, and δ = 0.1, constraining the parameter regression accuracy, intent matching degree, scene adaptability, and rule compliance of the multi-parameter collaborative decision-making intelligent network model for photography, respectively. Specifically, Lparam has the heaviest weight of 0.4, indicating the size of the gap between the camera adjustment parameters and the professional optimal value; Lintent has a weight of 0.3, indicating whether the camera adjustment parameters meet the user's shooting intent; Lscene has a weight of 0.2, indicating whether the camera adjustment parameters are suitable for the current shooting scene; and Lconstraint has a weight of 0.1, indicating whether the camera adjustment parameters generally conform to the basic constraints.
[0095] In this embodiment, a multi-task weighted loss function is used as the training constraint standard. The multi-parameter collaborative decision-making intelligent network model for photography is pre-trained offline or online using massive amounts of photographic samples. This allows the model to learn the intrinsic relationships between shooting intent, scene, camera parameters, and hardware boundaries, master the logic of multi-parameter collaborative adjustment, and complete the model's capability building. Through the above model pre-training strategy, a multi-parameter collaborative decision-making intelligent network model for photography that adapts to user preferences and professional photography adjustment schemes can be obtained.
[0096] In one embodiment, S30: the step of generating a camera adjustment parameter set based on a photographic multi-parameter collaborative decision-making intelligent network model further includes: Each of the aforementioned scene types is configured with corresponding camera parameter priority sorting and camera parameter hard constraint rules. The photography multi-parameter collaborative decision-making intelligent network model generates and optimizes the camera adjustment parameter set based on the camera parameter priority sorting and the camera parameter hard constraint rules.
[0097] For example, the intelligent network model for multi-parameter collaborative decision-making in photography incorporates scene-specific parameter adjustment weight matrices and insurmountable hard constraint rules. Higher weights result in higher adjustment priorities, fundamentally solving the image quality / effect loss problem caused by single-parameter adjustments. In this embodiment, corresponding camera parameter priority ranking and camera parameter hard constraint rules are set for each of the first preset number of scene types.
[0098] For example, the primary portrait scene category includes secondary categories such as outdoor front-lit portraits, outdoor backlit portraits, indoor portraits, low-light portraits, ID photo portraits, and group portraits. This scene classification covers the differences in lighting, environment, and number of people in portrait photography, adapting to different focus, exposure, and beautification logics. Specifically, the parameter priority order for portrait scenes, from highest to lowest, is: focus position, shutter speed, ISO sensitivity, aperture, white balance, and ISP parameters. Unbreakable hard constraints include: when there is no explicit need for depth-of-field adjustment, the aperture adjustment range must not exceed ±0.5EV; drastically reducing the aperture to adjust exposure, resulting in the disappearance of background blur, is prohibited; the shutter speed must not be lower than the safe shutter speed (1 divided by the equivalent focal length); and the ISO must not exceed 80% of the camera's native maximum ISO. For example, the primary landscape scene category includes secondary grand landscape, sunrise / sunset landscape, snow scene, rain / fog landscape, grassland / desert landscape, and seascape. This scene category encompasses landscape shooting scenarios with different lighting, weather, and terrain conditions, matching corresponding aperture, focus, and dynamic range adjustment logic. Specifically, the parameter priority order for landscape scenes, from highest to lowest, is: aperture, shutter speed, ISO sensitivity, white balance, and ISP parameters. The inviolable hard constraints are: ISO is prioritized and locked at its native minimum value; unless the shutter speed is lower than 1 / 30s (the safe shutter speed for a tripod), increasing ISO is prohibited; aperture is prioritized and locked within the optimal image quality range of f / 8-f / 11 for the lens; and focus is prioritized and set to hyperfocal distance.
[0099] For example, the primary night scene category includes secondary categories such as city night scene, night portrait, night landscape, handheld night scene, tripod night scene, and night light show. This scene classification distinguishes different subjects and shooting methods, matching corresponding shutter speed, ISO, and noise reduction priority control logic. Specifically, the parameter priority order for night scene scenarios, from highest to lowest, is: shutter speed, noise reduction algorithm, ISO sensitivity, aperture, and ISP parameters. Unbreakable hard constraints include: in handheld mode, the shutter speed must not be lower than the safe shutter speed; in tripod mode, priority should be given to extending the shutter speed, and blindly increasing ISO is prohibited; ISO increases must be linked to noise reduction level increases; and excessively increasing the shutter speed to brighten the image is prohibited.
[0100] For example, the primary backlight scene category includes secondary backlight portraits, backlight landscapes, backlight silhouettes, and backlight macro. This scene classification covers various backlight shooting subjects, matching corresponding HDR (High Dynamic Range) brightness optimization, highlight suppression, and subject exposure lock adjustment logic. Specifically, the parameter priority order for backlight scenes, from high to low, is: dynamic range optimization, focus position, exposure compensation, ISO shutter speed, and ISP parameters. The inviolable hard constraints are: prioritizing the HDR algorithm; locking focus and exposure to the core subject, preventing overexposure of the subject due to background brightening; and linking highlight suppression and shadow brightening algorithms.
[0101] For example, the first-level indoor scene includes the second-level indoor low-light scene, indoor flash scene, indoor still life scene, and indoor activity scene. This scene classification distinguishes different indoor lighting conditions and shooting objects, and matches the corresponding white balance, ISO sensitivity, and image stabilization control logic.
[0102] For example, the primary sports scene category includes secondary categories such as outdoor high-speed sports, indoor sports, pet sports, children's sports, and handheld autofocus tracking. This scene classification distinguishes different sports speeds and subjects, matching corresponding shutter speeds and focus mode priority control logic. Specifically, the parameter priority order for sports scenes, from highest to lowest, is: shutter speed, focus mode, ISO sensitivity, aperture, and ISP parameters. The inviolable hard constraints are: shutter speed must not be lower than 1 / 500s for regular sports scenes and not lower than 1 / 2000s for high-speed sports scenes; the focus mode is forcibly locked as continuous autofocus (AF-C); ISO restrictions can be relaxed, prioritizing ensuring shutter speed meets the requirements.
[0103] For example, the primary macro scene includes secondary flower macro, insect macro, still life macro, handheld macro, and tripod macro. This scene classification distinguishes the macro shooting subject and shooting support method, and matches the corresponding aperture, focus, and image stabilization control logic.
[0104] For example, the first-level pet scene includes the second-level pet static, pet dynamic, indoor pet, and outdoor pet. By combining the different states of the pet and the shooting environment through this scene classification, the corresponding focus mode and shutter speed control logic are matched.
[0105] For example, the primary starry sky scene includes secondary starry sky Milky Way, star trail photography, starry sky portrait, and night scene starry sky. By classifying the scene, different starry sky shooting targets and shooting methods are distinguished, and corresponding long exposure, high sensitivity and noise reduction control logic is matched.
[0106] For example, the first-level urban architectural scene includes second-level daytime architecture, nighttime architecture, architectural close-ups, and urban street photography. This scene classification distinguishes the lighting and shooting angle of architectural photography and matches the corresponding wide-angle correction, dynamic range, and sharpening adjustment logic.
[0107] In this embodiment, through multi-dimensional scene quantification perception and built-in scene-based hard constraint rules, the parameter adjustment is ensured to fully comply with the photography logic, avoiding problems such as blur disappearance, subject blurring, and noise explosion caused by single parameter adjustment in existing solutions. It has strong adaptability to all scenes and fundamentally avoids image quality loss.
[0108] In the above embodiments, an intelligent network model for multi-parameter collaborative decision-making in photography is built and optimized step by step through the offline training stage, the real-time data input stage, and the rule training constraint stage. This ensures that the camera adjustment parameters output by the intelligent network model not only meet the user's intent and the intelligent linkage requirements of the shooting scene, but also strictly maintain the bottom line of image quality and device performance, while taking into account the intelligence, matching degree, and imaging stability of parameter adjustment.
[0109] In one embodiment, S40: After the step of configuring the camera parameters according to the camera adjustment parameter set and synchronously updating the camera preview image, the following steps are included: Obtain user feedback information, and train and iteratively optimize the photography multi-parameter collaborative decision-making intelligent network model based on the feedback information.
[0110] For example, user feedback information is obtained, including secondary intent correction commands, shooting confirmation operations, and manual parameter adjustment operations. This feedback information, corresponding shooting intent features, scene awareness features, and parameter adjustment data are used to construct incremental training samples for local incremental training of the PMPCD-Net model, executed only when the camera is charging and idle. In this embodiment, during incremental training, 10% of the core samples from the pre-training dataset are retained to avoid catastrophic forgetting, and the deployed model is updated only when the model accuracy improves by at least 2%.
[0111] In this embodiment, through local incremental training, the PMPCD-Net model continuously adapts to the user's shooting preferences and aesthetic habits, achieving personalized adjustments for each user. The longer the user uses the model, the higher the degree of matching with their needs, thus realizing personalized self-learning iterative optimization of the PMPCD-Net model to adapt to user habits.
[0112] In the above embodiments, all models can be deployed locally offline, are 100% functionally available in network-free scenarios, have no privacy leakage risks, and are compatible with all imaging devices with ISP and NPU, making them highly adaptable.
[0113] Second Embodiment Based on the technical concept of the first embodiment of this application, this application also provides a camera parameter adjustment system, including: The shooting intention input module is used to obtain the shooting intention command input by the user's voice and convert the shooting intention command in voice form into shooting intention command in text form.
[0114] The intent parsing module is used to parse the shooting intent command and generate a shooting intent feature vector.
[0115] For example, the intent parsing module has a built-in NLU model that is finely tuned specifically for the photography domain.
[0116] The scene perception module is used to acquire camera preview images, identify and quantize the camera preview images, and generate scene perception feature vectors.
[0117] For example, the scene perception module has a built-in lightweight instance segmentation and scene classification dual-branch model of computer vision model optimized specifically for photography scenes.
[0118] The multi-parameter collaborative decision-making intelligent module is used to collect the camera's initial parameter feature vector and the camera's hardware capability boundary feature vector, and combine them with the shooting intention feature vector and the scene perception feature vector to generate a set of camera adjustment parameters based on the photography multi-parameter collaborative decision-making intelligent network model.
[0119] The parameter execution module is used to configure the camera parameters according to the camera adjustment parameter set and to update the camera preview screen synchronously.
[0120] For example, the parameter execution module includes a camera main control unit, a lens driving unit, an ISP image processing unit, and an image sensor.
[0121] The self-learning optimization module is used to obtain user feedback information and train and iteratively optimize the photography multi-parameter collaborative decision-making intelligent network model based on the feedback information.
[0122] For example, the self-learning optimization module completes the local incremental training and personalized iteration of the intelligent network model for multi-parameter collaborative decision-making in photography.
[0123] For example, the camera parameter adjustment system of this application is applicable to any smart terminal or imaging device capable of taking pictures, such as digital cameras, professional mirrorless cameras, professional SLR cameras, smartphones, action cameras, drone cameras, etc., and this application does not limit it.
[0124] In this embodiment, the camera system parameter adjustment has the following advantages compared to the prior art: 1. Completely eliminate the professional barrier to photography: Users do not need to master any professional photography knowledge. They can obtain professional-grade parameter solutions suitable for the scene simply by talking about their desired effect. According to actual tests, the excellent rate of non-professional users' photos has increased from 11% to 89%, fully releasing the camera hardware performance.
[0125] 2. Completely solve the core pain point of the disconnect between needs and parameters: Through photography-specific intent analysis technology, it achieves accurate conversion of spoken effect needs into structured shooting intentions. The accuracy rate of matching spoken needs is no less than 93%, and the accuracy rate of matching complex needs is no less than 86%, which is completely different from the shortcomings of existing solutions that can only recognize fixed parameter instructions.
[0126] 3. Extremely strong scene adaptability, fundamentally avoiding image quality loss: Through multi-dimensional scene quantitative perception and built-in scene-based hard constraint rules, it ensures that parameter adjustments fully comply with photographic logic, avoiding problems such as blurring, subject blurring, and noise explosion caused by single parameter adjustments in existing solutions, and adapting to all scenes.
[0127] 4. All models are deployed locally offline and are compatible with all devices: All models are deployed locally, ensuring 100% functionality even in network-free scenarios and eliminating the risk of privacy leaks; they are compatible with all imaging devices with ISP and NPU, making them highly adaptable.
[0128] 5. Self-learning and iteration, personalized adaptation to user habits: Through local incremental training, the model continuously adapts to users' shooting preferences and aesthetic habits, achieving personalized adjustments for each user. The longer the user uses the model, the higher the degree of matching with their needs.
[0129] Thirdly, this application also provides a camera, including: a processor and a memory storing a computer program, wherein when the processor runs the computer program, it implements the steps of the camera parameter adjustment method as described in the first embodiment.
[0130] For example, the camera includes at least a voice input unit, an image sensor, a processor, a memory, a lens driving unit, an ISP image processing unit, and an NPU neural network processor. The memory stores a computer program, and the processor and the NPU neural network processor execute the computer program to implement the camera parameter adjustment method of the first embodiment.
[0131] In conjunction with the first embodiment, this application also provides a specific implementation flow of a camera parameter adjustment method: Example 1: Intelligent adjustment process for exposure control requirements such as "too bright": 1. Shooting Intent Acquisition: The user issues a spoken command "Too bright", which is then converted into a textual shooting intent statement "Too bright" through speech recognition.
[0132] 2. Shooting Intent Analysis: The NLU model analyzes and outputs a structured intent, with a primary category of exposure control and a tertiary category of image overexposure adjustment, with an adjustment range of -2 and no constraints, outputting a 64-dimensional intent feature vector.
[0133] 3. Real-time scene perception: The CV model recognizes and outputs scene results. The shooting scene is an outdoor backlit portrait (confidence level 98.9%), with an ambient EV value of 14.5, a highlight ratio of 31%, a dynamic range of 13EV, and a color temperature of 5600K. The subject is a single face in the center of the image, with the face accounting for 24% of the image. The focus position is the face, and a 96-dimensional scene feature vector is output.
[0134] 4. AI Multi-Parameter Collaborative Decision Making: Input initial parameter set (aperture f / 1.8, shutter speed 1 / 1000s, ISO 200, AWB 5500K, exposure compensation 0EV) and hardware constraint features; the model, based on the hard constraints of portrait scenes, keeps the aperture f / 1.8 unchanged and outputs the target camera adjustment parameter set: ISO is reduced to the native minimum of 100, shutter speed is increased to 1 / 2000s, exposure compensation is reduced by 0.7EV, AWB remains unchanged, and the highlight suppression algorithm is enabled (intensity 65%).
[0135] 5. Parameter execution and preview update: The main controller synchronizes the target camera's parameter set with the VSYNC signal, all parameters take effect in the same frame, the preview update delay is no more than 80ms, the overexposure problem is solved, and the background blur effect is fully preserved.
[0136] 6. Iterative optimization: When the user issues a second correction command "It's still a bit bright", the AI model inherits the exposure adjustment intention, accumulates the adjustment range to -3 stops, further reduces the output exposure compensation by 0.3EV, increases the shutter speed to 1 / 3200s, and stores this sample in the incremental training set.
[0137] Example 2: Intelligent adjustment process for scene effects requirement "I want to shoot Tyndall rays": 1. Shooting Intent Acquisition: The user issues a spoken command, "I want to shoot Tyndall effect," which is then converted into a textual statement of shooting intent through speech recognition.
[0138] 2. Shooting Intent Analysis: The NLU model analyzes and outputs a structured intent: the first-level category is scene effects, the third-level sub-item is Tyndall effect enhancement, the adjustment range is +2, the constraint is to highlight the beam path and enhance the light and shadow layers, and the output intent feature vector is.
[0139] 3. Real-time scene perception: The CV model identifies and outputs scene results: The shooting scene is a backlit forest landscape (confidence level 97.8%), with side backlighting, ambient EV value 11.8, color temperature 5400K, the core subject is the beam path, the current focus position is the foreground trees, and the scene feature vector is output.
[0140] 4. AI Multi-Parameter Collaborative Decision-Making: Input initial parameter set (aperture f / 4, shutter speed 1 / 125s, ISO 100, contrast 0, special effects algorithm off); the model, based on the hard constraints of the landscape scene, keeps the native minimum ISO at 100, adjusts the aperture to f / 10, and outputs the target camera adjustment parameter set: aperture reduced to f / 10, shutter speed adjusted to 1 / 60s, exposure compensation reduced by 0.7EV, contrast increased by 25%, saturation increased by 15%, AI beam enhancement algorithm enabled (intensity 80%), and focus position adjusted to the center area of the beam.
[0141] 5. Parameter execution and preview update: All parameters take effect synchronously, and the Tyndall beam is clearly prominent in the preview screen with distinct light and shadow layers, perfectly matching user needs.
[0142] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0143] In this document, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, which includes not only the elements listed but also other elements not expressly listed.
[0144] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for adjusting camera parameters, characterized in that, include: In response to the shooting intention command, the shooting intention command is parsed to generate a shooting intention feature vector; Acquire a camera preview image, identify and quantize the camera preview image, and generate a scene-aware feature vector; Collect the camera initial parameter feature vector and the camera hardware capability boundary feature vector, combine them with the shooting intention feature vector and the scene perception feature vector, and generate a set of camera adjustment parameters based on the photography multi-parameter collaborative decision-making intelligent network model. Configure the camera parameters according to the camera adjustment parameter set, and update the camera preview screen synchronously.
2. The camera parameter adjustment method according to claim 1, characterized in that, Prior to the step of responding to the shooting intention command, the following are included: Obtain the user's voice input shooting intention command, and convert the voice shooting intention command into a text shooting intention command.
3. The camera parameter adjustment method according to claim 1, characterized in that, The step of parsing the shooting intention command and generating a shooting intention feature vector includes: The pre-trained natural language understanding model parses the shooting intention command and extracts three types of structured information: intention label, adjustment range, and constraint conditions. The structured information is encoded to generate a shooting intent feature vector that includes intent label encoding, adjustment amplitude value, and constraint condition encoding.
4. The camera parameter adjustment method according to claim 3, characterized in that, The steps of parsing the shooting intention command and extracting three types of structured information—intention label, adjustment range, and constraint conditions—based on the pre-trained natural language understanding model include: A four-level shooting intent classification system is constructed, with the levels subdivided from top to bottom: first-level classification, second-level classification, third-level classification, and fourth-level classification. The shooting intent command is matched hierarchically from top to bottom against the four-level shooting intent classification system to determine the classification category with the highest matching degree, so as to generate the intent tag.
5. The camera parameter adjustment method according to claim 4, characterized in that, The step of generating intent tags is followed by: The intent confidence value of the intent label is obtained based on the principle of classification probability normalization; When the intent confidence value is lower than the preset intent confidence threshold, scene information is added to fuse the shooting intent command, and the intent tag is regenerated based on the updated shooting intent command.
6. The camera parameter adjustment method according to claim 1, characterized in that, The steps of acquiring a camera preview image, identifying and quantizing the camera preview image, and generating a scene-aware feature vector include: The camera preview image is identified and quantified based on a pre-trained computer vision model, and scene information is extracted, including scene labels, environmental imaging parameters, and the subject being photographed. The scene information is encoded to generate a scene perception feature vector containing scene label encoding, environmental imaging parameter encoding, and subject encoding.
7. The camera parameter adjustment method according to claim 6, characterized in that, The steps of identifying and quantifying the camera preview image based on the pre-trained computer vision model, and extracting scene information, including scene labels, environmental imaging parameters, and the subject being photographed, include: Construct a first preset number of scene types, match the camera preview with each scene type, identify the scene type with the highest matching degree, and generate the scene tag; The pixel data in the camera preview image is statistically analyzed and quantified to generate the environmental imaging parameter information; A second preset number of subject feature types are constructed, and the camera preview image is matched with each subject feature type one by one to determine the corresponding feature value in order to generate the subject to be photographed.
8. The camera parameter adjustment method according to claim 7, characterized in that, The step of constructing a first preset number of scene types, matching the camera preview image with each scene type, and identifying the scene type with the highest matching degree to generate the scene label further includes: The scene confidence value for each scene type is obtained based on the principle of classification probability normalization. When the scene confidence value of any scene type is not lower than the preset scene confidence threshold, the scene type is confirmed as the scene type with the highest matching degree, and the scene label is generated.
9. The camera parameter adjustment method according to claim 1, characterized in that, The photographic multi-parameter collaborative decision-making intelligent network model includes an input layer, a fusion layer, a feature encoding layer, and an output layer; The step of collecting the camera initial parameter feature vector and the camera hardware capability boundary feature vector, combining them with the shooting intention feature vector and the scene perception feature vector, and generating a camera adjustment parameter set based on a multi-parameter collaborative decision-making intelligent network model for photography includes: The input layer inputs the camera initial parameter feature vector, the camera hardware capability boundary feature vector, the shooting intention feature vector, and the scene perception feature vector; The fusion layer fuses the shooting intent feature vector and the scene awareness feature vector, enabling the camera adjustment parameter set to associate the shooting intent command with the camera preview screen; The feature encoding layer is based on a transform encoder structure and learns the linkage logic in the initial parameters of the camera to adjust the camera adjustment parameter set. The output layer constrains the camera hardware capability boundary feature vector and outputs the camera adjustment parameter set.
10. The camera parameter adjustment method according to claim 9, characterized in that, Before the step of generating the camera adjustment parameter set based on the intelligent network model for multi-parameter collaborative decision-making in photography, the following steps are included: Construct a pre-trained dataset to train the photographic multi-parameter collaborative decision-making intelligent network model; The multi-task weighted evaluation rule is adopted to constrain and optimize the photography multi-parameter collaborative decision-making intelligent network model from four dimensions: parameter accuracy, intent matching degree, scene adaptability, and rule compliance.
11. The camera parameter adjustment method according to claim 8, characterized in that, The step of generating the camera adjustment parameter set based on the intelligent network model for multi-parameter collaborative decision-making in photography also includes: Each of the aforementioned scene types is configured with corresponding camera parameter priority sorting and camera parameter hard constraint rules. The photography multi-parameter collaborative decision-making intelligent network model generates and optimizes the camera adjustment parameter set based on the camera parameter priority sorting and the camera parameter hard constraint rules.
12. The camera parameter adjustment method according to any one of claims 1-11, characterized in that, The step of configuring the camera parameters according to the camera adjustment parameter set and synchronously updating the camera preview image includes: Obtain user feedback information, and train and iteratively optimize the photography multi-parameter collaborative decision-making intelligent network model based on the feedback information.
13. A camera parameter adjustment system, characterized in that, include: The shooting intention input module is used to obtain the user's voice input shooting intention command and convert the voice shooting intention command into a text shooting intention command; The intent parsing module is used to parse the shooting intent command and generate a shooting intent feature vector; The scene perception module is used to acquire camera preview images, identify and quantize the camera preview images, and generate scene perception feature vectors. The multi-parameter collaborative decision-making intelligent module is used to collect the camera initial parameter feature vector and the camera hardware capability boundary feature vector, and combine the shooting intention feature vector and the scene perception feature vector to generate a set of camera adjustment parameters based on the photography multi-parameter collaborative decision-making intelligent network model. The parameter execution module is used to configure the camera parameters according to the camera adjustment parameter set and to update the camera preview screen synchronously. The self-learning optimization module is used to obtain user feedback information and train and iteratively optimize the photography multi-parameter collaborative decision-making intelligent network model based on the feedback information.
14. A camera, characterized in that, include: A processor and a memory storing a computer program, wherein, when the processor runs the computer program, the steps of the camera parameter adjustment method according to any one of claims 1 to 12 are implemented.