Shooting method and device, computer equipment and computer readable storage medium
By generating shooting parameters through image recognition and processing models, the problem of poor shooting results and complex operation of smart devices in complex scenes is solved, and automated, user-friendly shooting is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUIZHOU TCL CLOUD INTERNET CORP TECH CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-04-24
AI Technical Summary
The automated shooting functions of existing smart devices cannot cope with complex, changing or mixed scenes, and cannot understand the user's abstract creative intent, resulting in unsatisfactory shooting results and complicated operation.
By acquiring preview images, image recognition is performed to obtain key elements, and a processing model is used to analyze and reason about the key elements to generate shooting parameters that meet the needs of the scene, and then shooting is performed in conjunction with user instructions.
It enables the automatic generation of shooting parameters that meet user needs in complex and mixed scenes, simplifying the operation process and improving shooting results.
Smart Images

Figure CN121924366A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, specifically to a shooting method, apparatus, computer equipment, and computer-readable storage medium. Background Technology
[0002] Automated shooting functions (such as automatic mode and scene recognition mode) are now very common in smartphones and digital cameras. These technologies are usually based on a preset rule base, using sensors to detect ambient light, identify a limited number of scenes (such as portraits, night scenes, and food), and then call preset parameter combinations to take pictures.
[0003] However, existing solutions typically make decisions based on fixed rules, making them unable to handle complex, variable, or mixed scenarios (such as the combination of backlit portraits and urban night scenes), and their level of intelligence is limited. At the same time, they cannot understand the user's abstract and subjective creative intentions (such as "to capture a melancholic film feel" or "to highlight the subject and blur the cluttered background").
[0004] This results in unsatisfactory shooting effects when the scene exceeds the preset rule library, and it cannot meet the user's personalized needs. At this time, the user still needs to manually enter the professional mode to make complex parameter adjustments, which is complicated. Summary of the Invention
[0005] This application provides a shooting method, apparatus, computer device, and computer-readable storage medium for automatically generating shooting parameters based on the current scene and using the shooting parameters to take a picture, thereby capturing a photo that meets the scene requirements.
[0006] The technical solution adopted by this invention to solve the problem is as follows: In a first aspect, this application provides a shooting method, comprising: acquiring a first preview image; performing image recognition on the first preview image to obtain key elements of the first preview image, the key elements being used to characterize image information of the first preview image; analyzing and reasoning about the key elements based on a processing model to obtain shooting parameters; and shooting an image based on the shooting parameters. In some embodiments of this application, before performing analysis and reasoning based on the key element to obtain the shooting parameters using a processing model, the method further includes: The system acquires shooting instructions, which include voice instructions, gesture instructions, text instructions, and / or control instructions. The control instructions are generated in response to the operation of relevant controls on the shooting device. The shooting instructions serve as input to the processing model, enabling the processing model to analyze and reason about the shooting instructions to obtain the corresponding shooting parameters.
[0007] In some embodiments of this application, the processing model analyzes and infers the key element to obtain the following imaging parameters: Based on this processing model, the key element and the shooting command are analyzed and reasoned to obtain the shooting parameters.
[0008] In some embodiments of this application, the image capture based on the shooting parameters includes: The shooting parameters are mapped to obtain the configuration parameters of the shooting device; Call the control interface of the shooting device to configure the configuration parameters for the shooting device; In response to the shooting operation, the image is captured based on the configuration parameters.
[0009] In some embodiments of this application, the parameter mapping of the shooting parameters to obtain the configuration parameters of the shooting device includes: The shooting parameters are mapped to obtain at least one set of configuration parameters for the shooting device; Display a first interface, which includes at least one set of configuration parameters and at least one option control, the at least one option control being used to represent the selection of configuration parameters; In response to an operation on the target option control, the target configuration parameter corresponding to the target option control is determined to be the configuration parameter of the shooting device.
[0010] In some embodiments of this application, before capturing an image based on the capturing parameters, after analyzing and reasoning about the key element based on a processing model to obtain the capturing parameters, the method further includes: The second interface is displayed, which includes a confirmation control used to indicate confirmation of the application of the shooting parameters; In response to the operation of the confirmation control, an operation to capture an image based on the shooting parameters is triggered.
[0011] In some embodiments of this application, the method further includes: A second interface is displayed, which includes a second preview image generated based on the shooting parameters.
[0012] Secondly, this application provides a shooting device, comprising: The acquisition module is used to acquire the first preview image; The processing module is used to perform image recognition on the first preview image to obtain key elements of the first preview image, which are used to characterize the image information of the first preview image; and to analyze and reason about the key elements based on the processing model to obtain the shooting parameters. The shooting module is used to capture images based on these shooting parameters.
[0013] Thirdly, this application also provides a computer device, which includes: One or more processors; Memory; and One or more applications, wherein the applications are stored in memory and configured to be executed by a processor to implement the focus recognition method of any of the first aspects.
[0014] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps in the focus identification method of any one of the first aspects.
[0015] The beneficial effects of this invention are as follows: During the shooting process, image recognition is performed on the preview image of the current scene captured by the camera to obtain the key elements of the current scene; then, a processing model is invoked to analyze and reason about these key elements to obtain shooting parameters that conform to the current scene; finally, the image is captured based on these shooting parameters. This automated provision of shooting parameters using a processing model allows users to capture photos that meet the requirements of the scene without requiring specialized knowledge, thus making the shooting solution applicable to various complex and mixed scenes. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a schematic diagram of the system architecture provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the module flow provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of an embodiment of the shooting method provided by the present invention; Figure 4a This is a schematic diagram of an interface provided in an embodiment of the present invention; Figure 4b This is another schematic diagram of an interface provided in an embodiment of the present invention; Figure 4c This is another schematic diagram of an interface provided in an embodiment of the present invention; Figure 4d This is another schematic diagram of an interface provided in an embodiment of the present invention; Figure 4e This is another schematic diagram of an interface provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of another embodiment of the shooting method provided in this invention; Figure 6 A schematic block diagram of a specific embodiment of the shooting device provided in this invention; Figure 7 This is a schematic block diagram of another specific embodiment of the shooting device provided in this invention; Figure 8 This is a schematic diagram of an embodiment of the computer device provided in this invention. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0019] In the description of this application, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," "third," etc., may explicitly or implicitly include one or more features.
[0020] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0021] It should be noted that since the method in this application embodiment is executed in a computer device, the processing objects of each computer device exist in the form of data or information, such as time, which is essentially time information. It is understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, they are all corresponding data that exist so that the computer device can process them. Specific details will not be elaborated here.
[0022] Automated shooting functions (such as automatic mode and scene recognition mode) are now widespread in smartphones and digital cameras. These technologies are typically based on a preset rule base, using sensors to detect ambient light, identify a limited number of scenes (such as portraits, night scenes, and food), and then call upon preset parameter combinations to take the picture. These decisions are usually based on fixed rules and cannot handle complex, variable, or mixed scenes (such as a combination of backlit portraits and urban night scenes), thus limiting their level of intelligence. At the same time, they cannot understand the user's abstract and subjective creative intentions (such as "to capture a melancholic film-like feel" or "to highlight the subject and blur the cluttered background"). As a result, when the scene exceeds the preset rule base, the shooting effect is often unsatisfactory and cannot meet the user's personalized needs. In this case, the user still needs to manually enter professional mode to adjust complex parameters, making the operation cumbersome.
[0023] To address this technical problem, this application provides the following technical solution: acquiring a first preview image; performing image recognition on the first preview image to obtain key elements of the first preview image, which are used to characterize the image information of the first preview image; analyzing and reasoning about the key elements based on a processing model to obtain shooting parameters; and capturing an image based on the shooting parameters. In this way, during the shooting process, the camera performs image recognition on the preview image of the current scene to obtain key elements of the current scene; then, the processing model is invoked to analyze and reason about the key elements to obtain shooting parameters that conform to the current scene; finally, an image is captured based on the shooting parameters. This automated provision of shooting parameters allows users to capture photos that meet the requirements of the scene without professional knowledge, thus making the shooting solution applicable to various complex and mixed scenes.
[0024] This application provides a shooting method, apparatus, computer device, and computer-readable storage medium for automatically generating shooting parameters based on the current scene and using these parameters to take a picture, thereby capturing a photo that meets the scene requirements. The electronic device provided in this application can be implemented as various types of user terminals or as a server.
[0025] By running the shooting method provided in the embodiments of this application, the electronic device can automatically generate shooting parameters based on the current scene and use the shooting parameters to take a picture, thereby taking a picture that meets the scene requirements.
[0026] The above methods can be applied to all smart devices with built-in cameras (such as mobile phones, tablets, drones, virtual reality (VR) devices or augmented reality (AR) devices), professional digital cameras, and even roadside cameras (such as optimizing image quality in low light and identifying specific events).
[0027] In one exemplary solution, the shooting method can be applied to image capture on a mobile phone. The image recognition model and processing model can be built into the phone; alternatively, they can be deployed on a cloud server and transmit images and data to the phone via a network. For example, in a scenario where the phone is capturing an image, after the user launches the phone's camera application, a first preview image of the current scene can be displayed on the screen. This first preview image can be uploaded to the image recognition model on the cloud server. The image recognition model analyzes the first preview image and obtains its key elements. These key elements include the subject being captured, environmental information, and composition information in the first preview image. These key elements are input into the processing model, which analyzes and infers from them to obtain corresponding shooting parameters. The cloud server feeds back these shooting parameters to the phone. The phone generates a second preview image based on these shooting parameters and displays it to the user. Simultaneously, a confirmation control and a cancellation control can be displayed on the phone's screen. The confirmation control instructs the user to apply the shooting parameters, and the cancellation control instructs the user not to apply the shooting parameters. Once the user confirms that the second preview image meets the requirements, they can click the confirmation control to apply the shooting parameters. The user then clicks the shooting control to obtain an image that meets the requirements.
[0028] It should be understood that the above is only an exemplary application scenario of the shooting method, and there are many other possible application scenarios, which are not limited here.
[0029] The shooting method provided in this application embodiment is applied to, for example, Figure 1 The system architecture diagram shown is for your reference. Figure 1To support a shooting method, the terminal device 100 connects to a server 300 via a network 200, and the server 300 connects to a database 400. The network 200 can be a wide area network (WAN), a local area network (LAN), or a combination of both. The client for implementing the shooting scheme is deployed on the terminal device 100, or it can run on the terminal device 100 as a standalone application. The specific form of the client is not limited here. The server 300 involved in this application can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. The terminal device 100 can be a device that includes both receiving and transmitting hardware, i.e., a device with receiving and transmitting hardware capable of performing bidirectional communication over a bidirectional communication link. Such a device can include cellular or other communication devices with single-line displays, multi-line displays, or cellular or other communication devices without multi-line displays. The specific terminal device 100 can be a desktop terminal or a mobile terminal. Specifically, it can also be one of the following: Augmented Reality (AR) devices, mobile phones, tablets, laptops, in-vehicle devices, wearable devices, smart TVs, smart home appliances, aircraft, or intelligent voice interaction devices (it should be understood that the terminal device can have a built-in camera function or can connect to an external camera). The solution provided in this application can be completed by the terminal device 100 and the server 300 working together. The database 400, in short, can be considered an electronic filing cabinet—a place to store electronic files, where users can perform operations such as adding, querying, updating, and deleting data. A "database" is a collection of data stored together in a certain way, shared by multiple users, with minimal redundancy, and independent of applications. A Database Management System (DBMS) is a computer software system designed for managing databases, generally possessing basic functions such as storage, retrieval, security, and backup. Database management systems can be categorized based on the database model they support, such as relational or Extensible Markup Language (XML); or based on the type of computer they support, such as server clusters or mobile phones; or based on the query language they use, such as Structured Query Language (SQL) or XQuery; or based on performance priorities, such as maximum scale or highest operating speed; or other classification methods.Regardless of the classification method used, some DBMSs can cross categories, for example, simultaneously supporting multiple query languages. In this application, database 400 can be used to store data such as shooting parameters, image recognition models, and processing models.
[0030] Those skilled in the art will understand that Figure 1 The system architecture diagram shown is an exemplary system architecture of the present application and does not constitute a limitation on the system architecture of the present application. Other system architectures may include more advanced architectures. Figure 1 The number of more or fewer terminal devices or servers shown, for example Figure 1 Only one server is shown in the diagram. It is understood that the system architecture may also include one or more other terminal devices or servers, which are not limited here.
[0031] It should be noted that, Figure 1 The system architecture shown is an example. The servers and scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of servers and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0032] Based on the above system architecture, the following will be used as an example. Figure 2 The module system shown illustrates the process of the shooting method of this application.
[0033] like Figure 2 As shown, the modular system of this shooting method includes a user interaction layer, an intelligent decision-making layer, an execution control layer, and a hardware layer. The user interaction layer can be understood as the interface interaction of the shooting device (i.e., the terminal device). The intelligent decision-making layer can be deployed on the shooting device or on a cloud server; no specific limitation is made here. The execution control layer can be understood as the internal execution layer of the shooting device. The hardware layer can be understood as the hardware modules of the shooting device.
[0034] In this application, the user interaction layer may include the following functional modules: A real-time preview screen is used to display the first and second preview images in this application.
[0035] The voice input interface is used to receive voice commands from the user.
[0036] The text input interface is used to receive text commands from the user.
[0037] This intelligent decision-making layer may include the following functional modules: Image or content analysis models are used to perform image recognition or content recognition on the first preview image to obtain key elements.
[0038] A multimodal large model engine is used to comprehensively analyze and process the key element, the voice command, and the text command to obtain the shooting parameters.
[0039] The parameter inference module is used to map the shooting parameters into configuration parameters that the shooting device can understand. That is, it converts the natural language suggestions or standardized parameter suggestions generated by the above multimodal large model into specific control instructions that the shooting device hardware can recognize.
[0040] The execution control layer may include the following functional modules: The shooting equipment parameter control application interface is used to call the corresponding hardware modules.
[0041] The parameter mapping and setting module is used to set the configuration parameters obtained by mapping the shooting parameters to the shooting device.
[0042] This hardware layer may include the following functional modules: Image sensors and lenses are used to capture images of the current scene.
[0043] The processor is used to process the image accordingly based on the configuration parameters.
[0044] The following describes the shooting method provided in this application using specific embodiments, such as... Figure 3 The diagram shown is a flowchart of one embodiment of the shooting method in this application. The shooting method is described below using a terminal device as the execution subject, and it may include the following steps 301-304, as detailed below: 301. Obtain the first preview image.
[0045] In this embodiment, a camera application can be installed on the terminal device. When the user wants to take a photo, they can click the icon of the camera application. The terminal device will receive a launch command for the corresponding camera application and launch it. At this time, the terminal device will be in a state of waiting to capture an image, meaning that the terminal device can activate its camera component. While the terminal device is in this state, it can acquire a first preview image through the camera component. This first preview image can be an image acquired by the camera component and displayed on the terminal device's corresponding display screen (e.g., a mobile phone screen or an external display for a drone). In other words, the first preview image can be an image acquired by the camera component before the user presses the shutter button.
[0046] Optionally, after the camera application is launched, the device can determine the current configuration parameters of the camera application in real time based on the ambient brightness and the color of the light sources in the environment. These parameters include combinations such as exposure value, shutter speed, and white balance.
[0047] 302. Perform image recognition on the first preview image to obtain the key elements of the first preview image, which are used to characterize the image information of the first preview image.
[0048] After acquiring the first preview image, the terminal device can call an image recognition model to perform image recognition on the first preview image in order to obtain the key elements of the first preview image.
[0049] Optionally, the image recognition model can be built into the terminal device or deployed on a cloud server. In this embodiment, the image recognition model is built into the terminal device. In this case, the image recognition model can be a lightweight model or a lightweight version of the cloud model before being deployed on the terminal device.
[0050] This image recognition model can select the appropriate network architecture based on actual needs. For example, MobileNet, YOLOv5s / YOLOv8n, EfficientNet-Lite, or other corresponding image processing models.
[0051] In this embodiment, the key elements include the subject being photographed, environmental information, and composition information.
[0052] The subject of the photograph can be a person, pet, building, landscape, flower, document, etc. The image recognition model can identify the subject, as well as its quantity, expression, and posture.
[0053] The environmental information may include lighting conditions (such as front lighting, backlighting, side lighting, weak light, strong light), weather (such as sunny, cloudy, rainy, snowy, etc.), time (such as daytime, dusk, nighttime, 12 noon, etc.), scene type (such as indoor, outdoor, beach, snow, etc.).
[0054] The composition information includes the position of the subject, the clutter level of the background, the guiding lines, and the spatial hierarchy.
[0055] 303. Based on the processing model, analyze and reason about the key element to obtain the shooting parameters.
[0056] In this embodiment, after acquiring the key element, the terminal device inputs the key element into the processing model, so that the processing model analyzes and infers the key element to obtain the corresponding shooting parameters.
[0057] Optionally, to better meet the user's needs for photo quality, the terminal device can also receive the user's photo-taking command and input the command along with the key element into the processing model for comprehensive analysis and reasoning to obtain the final shooting parameters. An example of this approach is as follows: The terminal device acquires a shooting instruction, which includes voice instructions, text instructions, and / or control instructions; then, based on the processing model, it analyzes and infers the shooting instruction and the key element to obtain the shooting parameters.
[0058] The voice command can be obtained through a voice input interface. This voice input interface can be a microphone built into the terminal device or an external microphone.
[0059] The text instruction can be obtained through a text input interface. This text input interface can be a corresponding input box or a text upload control displayed on the terminal device.
[0060] The control command can be generated in response to physical controls (such as physical buttons) or virtual controls (such as controls displayed on the user interface of the terminal device). This control command can be used to represent the user's shooting intention. For example, by tapping the style selection control on the user interface of the terminal device, the user can select the corresponding shooting style.
[0061] Optionally, the processing model can directly output the shooting parameters. Alternatively, the processing model can output natural language suggestions, which are then further processed with the current configuration parameters of the camera app to obtain the shooting parameters. The specific approach is not limited here. For example, assuming that the image recognition model can identify the key elements including "backlight", "portrait", "dusk" and "beach", the processing model can deduce the following natural language suggestions: "increase exposure compensation to brighten the face", "appropriately reduce highlights to avoid overexposure of the sky", "warm color temperature to enhance the atmosphere"; then the natural language suggestions are generated and parameterized to obtain the shooting parameters (such as exposure value (EV+1.7), shutter speed (1 / 200s), sensitivity (ISO (200), white balance (5500K) and other parameter combinations).
[0062] For example, suppose the image recognition model can identify key elements including "still life", "cake", and "warm indoor light", and the user's voice input command is "take a bright and fresh photo". The processing model can understand "bright and fresh" in the voice command as follows: the photo should have a cool tone overall, and the photo should have a clean background. At the same time, the natural language suggestions obtained by the processing model in conjunction with the existing key elements and the voice command analysis can be as follows: significantly increase the exposure, shift the white balance to blue, and use a large aperture to further blur the background. Then, based on the natural language suggestions, parameterization is performed to obtain the corresponding shooting parameters (such as shutter speed (1 / 100s), ISO (300), aperture value (f / 1.2), white balance (3500K), etc.).
[0063] It should be understood that this processing model can be a pre-trained multimodal model based on a large language model. A large language model (LLM) is a deep learning-based artificial intelligence model trained on massive amounts of data, capable of understanding, generating, and processing natural language, and possessing powerful reasoning and knowledge integration capabilities. This multimodal processing model refers to a model capable of simultaneously processing multiple modalities of data, such as text, images, audio, and video. Its core is to achieve understanding, conversion, or fusion between different modalities through cross-modal association learning. Depending on the application scenario, this multimodal processing model can include general cross-modal understanding models (such as CLIP (Contrastive Language-Image Pretraining), ALBEF (Aligning Language and Vision with BERT), FLAVA (A Foundational Language And Vision Alignment Model), etc.), generative multimodal models (such as processing models that generate images based on text), video, text, and audio multimodal models, etc.
[0064] In this embodiment, to improve the inference effectiveness of the processing model, the processing model can adopt the following possible implementation methods during the training process: In one exemplary scheme, the processing model can learn from a massive amount of photographic theory texts, transforming abstract aesthetic principles into computable logic. For example, it covers basic aesthetics such as composition (rule of thirds, leading lines, framing), color matching (complementary colors, warm and cool contrast, tonal unity), and the use of light and shadow (differences in the effects of front lighting / backlighting / side lighting), as well as stylized aesthetics (such as the low saturation of Japanese fresh style, the high contrast of European and American documentary style, and the negative space composition of traditional Chinese style photography), and can distinguish the core characteristics of different styles.
[0065] In another exemplary scheme, the processing model can learn optics and camera hardware-related knowledge to master the underlying technical rules of shooting. Among them, the core principles of optics and camera hardware-related knowledge include the three elements of exposure (the linkage between aperture and shutter speed), focusing mechanism (applicable scenarios of single-point focus / continuous focus), depth of field control (the relationship between aperture size and degree of blur), and environmental adaptation logic (such as how to balance ISO and shutter speed in low light to avoid noise or blur; how to control exposure in strong light to avoid overexposure).
[0066] Exposure describes the total amount of light received by the film or image sensor.
[0067] Shutter speed refers to the length of time a camera shutter is open, affecting the amount of light entering the camera and motion blur.
[0068] ISO sensitivity refers to the sensitivity of film or image sensor to light, affecting image brightness and noise.
[0069] White balance is used to adjust the representation of white in an image to correct color deviations under different light sources.
[0070] Another exemplary approach involves using high-quality photographs with labels or scene descriptions as training samples to train the processing model. For example, it identifies high-frequency features in portrait photography: the subject is located at the intersection of the rule of thirds, the background blur is moderate, the face is accurately exposed, and the eyes are clear. Another example is summarizing the commonalities of excellent night scene photography: preserving details in shadows, controlling highlight clipping, and making reasonable use of lighting as accents.
[0071] It should be understood that the above three training methods can be used in combination to improve the analytical reasoning ability of the processing model.
[0072] Optionally, the processing model can be deployed on the terminal device or on a cloud server. In this embodiment, the processing model is deployed on the terminal device, and to adapt to the terminal device, the processing model can be a lightweight model or a pre-trained multimodal model that has been lightweighted.
[0073] 304. Capture images based on these shooting parameters.
[0074] In this embodiment, after the processing model outputs the shooting parameters, it needs to convert them into specific control instructions that can be recognized by the hardware of the terminal device's camera program; then, based on the control instructions, it configures the camera program's configuration parameters; and finally, in response to the user's camera operation, it takes a picture based on the configuration parameters.
[0075] In one exemplary solution, the terminal device can perform parameter mapping on the shooting parameters to obtain the configuration parameters of the shooting device (such as a mobile phone camera); call the control interface of the shooting device to configure the configuration parameters for the shooting device; and in response to the shooting operation, capture the image based on the configuration parameters. For example, when the shooting parameters are: exposure value (EV+1.7), shutter speed (1 / 200s), sensitivity (ISO) (200), and white balance (5500K), the shooting parameters are parameter mapped to generate configuration parameters that the camera can understand; then the operating system of the camera hardware (or mobile phone camera) will write the configuration parameters into the camera hardware through the official control interface (such as Android's Camera2 API, iOS's AVFoundation), so that the hardware can adjust according to the instructions (such as actually extending the sensor's exposure time). For example, on an Android phone, the Camera2 API is used to find the interface that controls the exposure time (CaptureRequest.SENSOR_EXPOSURE_TIME is the "exposure time parameter identifier" defined in this API); the value is passed in through API methods (such as set()), i.e., the code executes set(CaptureRequest.SENSOR_EXPOSURE_TIME, 200000L). It should be understood that 200000L is the configuration parameter, where L represents a long integer type (conforming to the format required by the API). After receiving the instruction from the API, the camera sensor actually adjusts the exposure time to 200000 microseconds (i.e., 0.2 seconds), and the final image becomes brighter.
[0076] Optionally, to improve the user's interactive experience, a solution can be provided that allows users to select application configuration parameters and / or shooting parameters.
[0077] In one exemplary solution, when the terminal device performs parameter mapping based on the shooting parameters, it can generate at least one set of configuration parameters. A first interface is displayed on the display device associated with the terminal device. This first interface can include the at least one set of configuration parameters and at least one option control. Each set of configuration parameters corresponds to an option control, which is used to select a configuration parameter. Selecting the option control confirms the application of its corresponding configuration parameter. In response to an operation on the target option control (such as a single click or long press), the target configuration parameter corresponding to the target option control is determined as the configuration parameter of the shooting device. Figure 4aAs shown, an exemplary scheme for displaying the first interface is provided, wherein the first interface includes configuration parameter combination 1 and configuration parameter combination 2, wherein configuration parameter combination 1 corresponds to option control 1 and configuration parameter combination 2 corresponds to option control 2; then, by clicking on option control 1, it is confirmed that configuration parameter combination 1 is selected as the configuration parameter of the shooting application or shooting device associated with the terminal device.
[0078] In another exemplary embodiment, after the terminal device acquires the shooting parameters, a second interface is displayed on the display device associated with the terminal device. This second interface includes a confirmation control and / or a cancellation control. The confirmation control indicates confirmation to apply the shooting parameters to the shooting application associated with the terminal device (i.e., a built-in shooting application with an integrated camera lens) or the shooting device (i.e., an external shooting device or camera lens, such as a drone controller with a display screen and a camera lens carried by the drone). The cancellation control indicates not to apply the shooting parameters to the shooting application or camera device associated with the terminal device. In response to the operation of the confirmation control, an operation to capture an image based on the shooting parameters is triggered. Figure 4b As shown, the second interface can display the shooting parameters and provide "confirm" and "cancel" controls.
[0079] Optional, in Figure 4b In the scenario shown, after the user clicks the "Confirm" control, the user can be redirected to a page like this. Figure 4a The interface shown will then prompt you to select the corresponding configuration parameters.
[0080] Optionally, to make it easier for users to view the image effects corresponding to the shooting parameters and to confirm whether to apply the shooting parameters based on these image effects, an example solution is as follows: The display device associated with the terminal device can display the second interface, which can also display a second preview image generated based on the shooting parameters. For example... Figure 4c As shown, the second interface can directly display the second preview image and indicate that the second preview image is generated based on the shooting parameters, and also includes the "Confirm" and "Cancel" controls.
[0081] It should be understood that, since the shooting parameters can generate at least one set of configuration parameters during the parameter mapping process, the terminal device can also display multiple second preview images, with each preview image corresponding to a set of configuration parameters. The second interface also includes a switching control, which indicates whether to switch between displaying different second preview images corresponding to different configuration parameters on the second interface. The second interface also includes a confirmation control, which indicates whether to confirm the application of the configuration parameters. Figure 4dAs shown, the second interface displays a second preview image, a toggle control, and a confirmation control. It should be understood that the toggle control can be a separate icon control, such as... Figure 4d As shown. This toggle control can also be an icon control used to represent this configuration parameter, such as... Figure 4e As shown.
[0082] It should be understood that the above solutions are merely exemplary and no specific limitations are made here.
[0083] The above description uses a terminal device as the execution entity to illustrate the shooting method. The following description uses a combination of a terminal device and a server to execute the shooting method, where the server deploys image recognition and processing models. For example... Figure 5 The diagram shown is a flowchart of one embodiment of the shooting method in this application. The shooting method can be described as follows: Steps 501 to 506 are as follows: 501. The terminal device acquires the first preview image.
[0084] In this embodiment, a camera application can be installed on the terminal device. When the user wants to take a photo, they can click the icon of the camera application. The terminal device will receive a launch command for the corresponding camera application and launch it. At this time, the terminal device will be in a state of waiting to capture an image, meaning that the terminal device can activate its camera component. While the terminal device is in this state, it can acquire a first preview image through the camera component. This first preview image can be an image acquired by the camera component and displayed on the terminal device's corresponding display screen (e.g., a mobile phone screen or an external display for a drone). In other words, the first preview image can be an image acquired by the camera component before the user presses the shutter button.
[0085] Optionally, after the camera application is launched, the device can determine the current configuration parameters of the camera application in real time based on the ambient brightness and the color of the light sources in the environment. These parameters include combinations such as exposure value, shutter speed, and white balance.
[0086] 502. The terminal device uploads the first preview image to the server.
[0087] In this embodiment, in order to protect user privacy, the terminal device encrypts the first preview image when uploading it to the server, and then uploads the encrypted first preview image to the server.
[0088] 503. The server performs image recognition on the first preview image to obtain the key elements of the first preview image, which are used to characterize the image information of the first preview image.
[0089] After obtaining the first preview image, the server can call an image recognition model to perform image recognition on the first preview image in order to obtain the key elements of the first preview image.
[0090] In this embodiment, the image recognition model is deployed on the server, and the image recognition model can be the full model.
[0091] The image recognition model can be selected based on the actual needs, choosing the appropriate network architecture. Examples include MobileNet, YOLOv5s / YOLOv8n, EfficientNet-Lite, or other suitable image processing models.
[0092] In this embodiment, the key elements include the subject being photographed, environmental information, and composition information.
[0093] The subject of the photograph can be a person, pet, building, landscape, flower, document, etc. The image recognition model can identify the subject, as well as its quantity, expression, and posture.
[0094] The environmental information may include lighting conditions (such as front lighting, backlighting, side lighting, weak light, strong light), weather (such as sunny, cloudy, rainy, snowy, etc.), time (such as daytime, dusk, nighttime, 12 noon, etc.), scene type (such as indoor, outdoor, beach, snow, etc.).
[0095] The composition information includes the position of the subject, the clutter level of the background, the guiding lines, and the spatial hierarchy.
[0096] 504. The server analyzes and infers the key element based on the processing model to obtain the shooting parameters.
[0097] In this embodiment, after obtaining the key element, the server inputs the key element into the processing model, so that the processing model analyzes and infers the key element to obtain the corresponding shooting parameters.
[0098] Optionally, in order to make the photo effect more in line with the user's needs, the server can also receive the photo shooting command uploaded by the terminal device, and input the photo shooting command and the key element into the processing model for comprehensive analysis and reasoning to obtain the final shooting parameters.
[0099] Optionally, to ensure user privacy, the terminal device may encrypt the photo-taking command when uploading it to the server.
[0100] In this embodiment, the shooting instruction includes voice instructions, text instructions, and / or control instructions; then, based on the processing model, the shooting instruction and the key element are analyzed and reasoned to obtain the shooting parameters.
[0101] The voice command can be obtained through the voice input interface of the terminal device. This voice input interface can be either a built-in microphone of the terminal device or an external microphone.
[0102] The text command can be obtained through the text input interface of the terminal device. This text input interface can be a corresponding input box or a text upload control displayed on the terminal device.
[0103] The control command can be generated in response to physical controls (such as physical buttons) or virtual controls (such as controls displayed on the user interface of the terminal device). This control command can be used to represent the user's shooting intention. For example, by tapping the style selection control on the user interface of the terminal device, the user can select the corresponding shooting style.
[0104] Optionally, the processing model can directly output the shooting parameters. Alternatively, the processing model can output natural language suggestions, and then further process these suggestions with the current configuration parameters of the camera app to obtain the shooting parameters. The specific approach is not limited here. For example, assuming that the image recognition model can identify the key elements including "backlight", "portrait", "dusk" and "beach", the processing model can deduce the following natural language suggestions: "increase exposure compensation to brighten the face", "appropriately reduce highlights to avoid overexposure of the sky", "warm color temperature to enhance the atmosphere"; then the natural language suggestions are generated and parameterized to obtain the shooting parameters (such as exposure value (EV+1.7), shutter speed (1 / 200s), sensitivity (ISO (200), white balance (5500K) and other parameter combinations).
[0105] For example, suppose the image recognition model can identify key elements including "still life", "cake", and "warm indoor light", and the user's voice input command is "take a bright and fresh photo". The processing model can understand "bright and fresh" in the voice command as follows: the photo should have a cool tone overall, and the photo should have a clean background. At the same time, the natural language suggestions obtained by the processing model in conjunction with the existing key elements and the voice command analysis can be as follows: significantly increase the exposure, shift the white balance to blue, and use a large aperture to further blur the background. Then, based on the natural language suggestions, parameterization is performed to obtain the corresponding shooting parameters (such as shutter speed (1 / 100s), ISO (300), aperture value (f / 1.2), white balance (3500K), etc.).
[0106] It should be understood that this processing model can be a pre-trained multimodal model based on a large language model. A large language model (LLM) is a deep learning-based artificial intelligence model trained on massive amounts of data, capable of understanding, generating, and processing natural language, and possessing powerful reasoning and knowledge integration capabilities. This multimodal processing model refers to a model capable of simultaneously processing multiple modalities of data, such as text, images, audio, and video. Its core is to achieve understanding, conversion, or fusion between different modalities through cross-modal association learning. Depending on the application scenario, this multimodal processing model can include general cross-modal understanding models (such as CLIP (Contrastive Language-Image Pretraining), ALBEF (Aligning Language and Vision with BERT), FLAVA (A Foundational Language And Vision Alignment Model), etc.), generative multimodal models (such as processing models that generate images based on text), video, text, and audio multimodal models, etc.
[0107] In this embodiment, to improve the inference effectiveness of the processing model, the processing model can adopt the following possible implementation methods during the training process: In one exemplary scheme, the processing model can learn from a massive amount of photographic theory texts, transforming abstract aesthetic principles into computable logic. For example, it covers basic aesthetics such as composition (rule of thirds, leading lines, framing), color matching (complementary colors, warm and cool contrast, tonal unity), and the use of light and shadow (differences in the effects of front lighting / backlighting / side lighting), as well as stylized aesthetics (such as the low saturation of Japanese fresh style, the high contrast of European and American documentary style, and the negative space composition of traditional Chinese style photography), and can distinguish the core characteristics of different styles.
[0108] In another exemplary scheme, the processing model can learn optics and camera hardware-related knowledge to master the underlying technical rules of shooting. Among them, the core principles of optics and camera hardware-related knowledge include the three elements of exposure (the linkage between aperture and shutter speed), focusing mechanism (applicable scenarios of single-point focus / continuous focus), depth of field control (the relationship between aperture size and degree of blur), and environmental adaptation logic (such as how to balance ISO and shutter speed in low light to avoid noise or blur; how to control exposure in strong light to avoid overexposure).
[0109] Exposure describes the total amount of light received by the film or image sensor.
[0110] Shutter speed refers to the length of time a camera shutter is open, affecting the amount of light entering the camera and motion blur.
[0111] ISO sensitivity refers to the sensitivity of film or image sensor to light, affecting image brightness and noise.
[0112] White balance is used to adjust the representation of white in an image to correct color deviations under different light sources.
[0113] Another exemplary approach involves using high-quality photographs with labels or scene descriptions as training samples to train the processing model. For example, it identifies high-frequency features in portrait photography: the subject is located at the intersection of the rule of thirds, the background blur is moderate, the face is accurately exposed, and the eyes are clear. Another example is summarizing the commonalities of excellent night scene photography: preserving details in shadows, controlling highlight clipping, and making reasonable use of lighting as accents.
[0114] It should be understood that the above three training methods can be used in combination to improve the analytical reasoning ability of the processing model.
[0115] 505. The server sends the shooting parameters to the terminal device.
[0116] In this embodiment, to ensure user privacy, the server can encrypt the shooting parameters before sending them to the terminal device.
[0117] 506. The terminal device captures images based on these shooting parameters.
[0118] In this embodiment, after receiving the shooting parameters, the terminal device needs to convert them into specific control instructions that can be recognized by the hardware of the terminal device's camera program; then, it configures the configuration parameters of the camera program based on the control instructions; and finally, in response to the user's camera operation, it takes a picture based on the configuration parameters.
[0119] In one exemplary solution, the terminal device can perform parameter mapping on the shooting parameters to obtain the configuration parameters of the shooting device (such as a mobile phone camera); call the control interface of the shooting device to configure the configuration parameters for the shooting device; and in response to the shooting operation, capture the image based on the configuration parameters. For example, when the shooting parameters are: exposure value (EV+1.7), shutter speed (1 / 200s), sensitivity (ISO) (200), and white balance (5500K), the shooting parameters are parameter mapped to generate configuration parameters that the camera can understand; then the operating system of the camera hardware (or mobile phone camera) will write the configuration parameters into the camera hardware through the official control interface (such as Android's Camera2 API, iOS's AVFoundation), so that the hardware can adjust according to the instructions (such as actually extending the sensor's exposure time). For example, on an Android phone, the Camera2 API is used to find the interface that controls the exposure time (CaptureRequest.SENSOR_EXPOSURE_TIME is the "exposure time parameter identifier" defined in this API); the value is passed in through API methods (such as set()), i.e., the code executes set(CaptureRequest.SENSOR_EXPOSURE_TIME, 200000L). It should be understood that 200000L is the configuration parameter, where L represents a long integer type (conforming to the format required by the API). After receiving the instruction from the API, the camera sensor actually adjusts the exposure time to 200000 microseconds (i.e., 0.2 seconds), and the final image becomes brighter.
[0120] Optionally, to improve the user's interactive experience, a solution can be provided that allows users to select application configuration parameters and / or shooting parameters.
[0121] In one exemplary solution, when the terminal device performs parameter mapping based on the shooting parameters, it can generate at least one set of configuration parameters. A first interface is displayed on the display device associated with the terminal device. This first interface can include the at least one set of configuration parameters and at least one option control. Each set of configuration parameters corresponds to an option control, which is used to select a configuration parameter. Selecting the option control confirms the application of its corresponding configuration parameter. In response to an operation on the target option control (such as a single click or long press), the target configuration parameter corresponding to the target option control is determined as the configuration parameter of the shooting device. Figure 4aAs shown, an exemplary scheme for displaying the first interface is provided, wherein the first interface includes configuration parameter combination 1 and configuration parameter combination 2, wherein configuration parameter combination 1 corresponds to option control 1 and configuration parameter combination 2 corresponds to option control 2; then, by clicking on option control 1, it is confirmed that configuration parameter combination 1 is selected as the configuration parameter of the shooting application or shooting device associated with the terminal device.
[0122] In another exemplary embodiment, after the terminal device acquires the shooting parameters, a second interface is displayed on the display device associated with the terminal device. This second interface includes a confirmation control and / or a cancellation control. The confirmation control indicates confirmation to apply the shooting parameters to the shooting application associated with the terminal device (i.e., a built-in shooting application with an integrated camera lens) or the shooting device (i.e., an external shooting device or camera lens, such as a drone controller with a display screen and a camera lens carried by the drone). The cancellation control indicates not to apply the shooting parameters to the shooting application or camera device associated with the terminal device. In response to the operation of the confirmation control, an operation to capture an image based on the shooting parameters is triggered. Figure 4b As shown, the second interface can display the shooting parameters and provide "confirm" and "cancel" controls.
[0123] Optionally, to make it easier for users to view the image effects corresponding to the shooting parameters and to confirm whether to apply the shooting parameters based on these image effects, an example solution is as follows: The display device associated with the terminal device can display the second interface, which can also display a second preview image generated based on the shooting parameters. For example... Figure 4c As shown, the second interface can directly display the second preview image and indicate that the second preview image is generated based on the shooting parameters, and also includes the "Confirm" and "Cancel" controls.
[0124] It should be understood that, since the shooting parameters can generate at least one set of configuration parameters during the parameter mapping process, the terminal device can also display multiple second preview images, with each preview image corresponding to a set of configuration parameters. The second interface also includes a switching control, which indicates whether to switch between displaying different second preview images corresponding to different configuration parameters on the second interface. The second interface also includes a confirmation control, which indicates whether to confirm the application of the configuration parameters. Figure 4d As shown, the second interface displays a second preview image, a toggle control, and a confirmation control. It should be understood that the toggle control can be a separate icon control, such as... Figure 4d As shown. This toggle control can also be an icon control used to represent this configuration parameter, such as... Figure 4e As shown.
[0125] It should be understood that the above solutions are merely exemplary and no specific limitations are made here.
[0126] To better implement the shooting method in the embodiments of this application, based on the shooting method, the embodiments of this application also provide a shooting device, such as... Figure 6 As shown, the imaging device 600 includes: Module 601 is used to acquire preview images; The processing module 602 is used to perform image recognition on the preview image to obtain key elements of the preview image, which are used to characterize the image information of the first preview image; and to analyze and reason about the key elements based on the processing model to obtain the shooting parameters. The imaging module 603 is used to capture images based on the imaging parameters.
[0127] In this embodiment, during the shooting process, image recognition is performed on the preview image of the current scene captured by the camera to obtain the key elements of the current scene; then, a processing model is invoked to analyze and reason about these key elements to obtain shooting parameters that conform to the current scene; finally, the image is captured based on these shooting parameters. This automated provision of shooting parameters allows users to capture photos that meet the requirements of the scene without professional knowledge, thus making the shooting solution applicable to various complex and mixed scenes.
[0128] In some embodiments of this application, such as Figure 6 As shown, the acquisition module 601 is also used for: The system acquires shooting instructions, which include voice instructions, gesture instructions, text instructions, and / or control instructions. The control instructions are generated in response to the operation of relevant controls on the shooting device. The shooting instructions serve as input to the processing model, enabling the processing model to analyze and reason about the shooting instructions.
[0129] In some embodiments of this application, such as Figure 6 As shown, the processing module 602 is specifically used for: Based on this processing model, the key element and the shooting command are analyzed and reasoned to obtain the shooting parameters.
[0130] In some embodiments of this application, such as Figure 6 As shown, the shooting module 603 is specifically used for: The shooting parameters are mapped to obtain the configuration parameters of the shooting device; Call the parameter control interface of the shooting device to configure the shooting device with the configuration parameters; In response to the shooting operation, the image is captured based on the configuration parameters.
[0131] In some embodiments of this application, such as Figure 7 As shown, the shooting module 603 is specifically used for: The shooting parameters are mapped to obtain at least one set of configuration parameters for the shooting device; The shooting device also includes a display module 604 for displaying a first interface, which includes at least one set of configuration parameters and at least one option control, the at least one option control being used to represent selecting configuration parameters; The processing module 602 is used to determine, in response to an operation on the target option control, that the target configuration parameter corresponding to the target option control is the configuration parameter of the shooting device.
[0132] In some embodiments of this application, such as Figure 7 As shown, the display module 604 is also used for: The second interface is displayed, which includes a confirmation control used to indicate confirmation of the application of the shooting parameters; In response to the operation of the confirmation control, an operation to capture an image based on the shooting parameters is triggered.
[0133] In some embodiments of this application, such as Figure 7 As shown, the display module 604 is also used for: The second interface is displayed, which includes a second preview image generated based on the shooting parameters.
[0134] This application also provides a computer device that integrates any of the shooting devices provided in this application. The computer device includes: One or more processors; Memory; and One or more applications, wherein the applications are stored in memory and configured to be executed by a processor from the steps of the shooting method in any of the embodiments described above.
[0135] This application also provides a computer device that integrates any of the shooting devices provided in this application. For example... Figure 8 As shown, it illustrates a structural schematic diagram of the computer device involved in the embodiments of this application, specifically: The computer device may include components such as a processor 801 with one or more processing cores, a memory 802 with one or more computer-readable storage media, a power supply 803, and an input unit 804. Those skilled in the art will understand that... Figure 8 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 801 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 802, and by calling data stored in the memory 802, thereby providing overall monitoring of the computer device. Optionally, the processor 801 may include one or more processing cores; preferably, the processor 801 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 801.
[0136] The memory 802 can be used to store software programs and modules. The processor 801 executes various functional applications and data processing by running the software programs and modules stored in the memory 802. The memory 802 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 802 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 802 may also include a memory controller to provide the processor 801 with access to the memory 802.
[0137] The computer device also includes a power supply 803 that supplies power to the various components. Preferably, the power supply 803 can be logically connected to the processor 801 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 803 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0138] The computer device may also include an input unit 804, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0139] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 801 in the computer device loads the executable files corresponding to the processes of one or more application programs into the memory 802 according to the following instructions, and the processor 801 runs the application programs stored in the memory 802 to realize various functions, as follows: Get the first preview image; Image recognition is performed on the first preview image to obtain the key elements of the first preview image, which are used to characterize the image information of the first preview image; The key element is analyzed and reasoned based on the processing model to obtain the shooting parameters; Images are captured based on these shooting parameters.
[0140] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0141] Therefore, embodiments of this application provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any of the cell handover methods provided in embodiments of this application. For example, the computer program loaded by the processor can execute the following steps: Get the first preview image; Image recognition is performed on the first preview image to obtain the key elements of the first preview image, which are used to characterize the image information of the first preview image; The key element is analyzed and reasoned based on the processing model to obtain the shooting parameters; Images are captured based on these shooting parameters.
[0142] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.
[0143] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.
[0144] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0145] The foregoing has provided a detailed description of a shooting method, apparatus, device, and storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A shooting method, characterized in that, include: Get the first preview image; Image recognition is performed on the first preview image to obtain key elements of the first preview image, and the key elements are used to characterize the image information of the first preview image; The key elements are analyzed and reasoned based on the processing model to obtain the shooting parameters; Images are captured based on the aforementioned shooting parameters.
2. The method according to claim 1, characterized in that, Before analyzing and reasoning about the key elements based on the processing model to obtain the shooting parameters, the method further includes: The system acquires shooting instructions, which include voice instructions, gesture instructions, text instructions, and / or control instructions. The control instructions are generated in response to operations on relevant controls of the shooting device. The shooting instructions serve as input to the processing model, enabling the processing model to analyze and reason about the shooting instructions.
3. The method according to claim 2, characterized in that, The analysis and reasoning of the key elements based on the processing model to obtain the shooting parameters includes: The processing model is used to analyze and reason about the key elements and the shooting instructions to obtain the shooting parameters.
4. The method according to any one of claims 1 to 3, characterized in that, The image capture based on the shooting parameters includes: The shooting parameters are mapped to obtain the configuration parameters of the shooting device; The control interface of the shooting device is invoked to configure the configuration parameters for the shooting device; In response to the shooting operation, the image is captured based on the configuration parameters.
5. The method according to claim 4, characterized in that, The step of mapping the shooting parameters to obtain the configuration parameters of the shooting device includes: The shooting parameters are mapped to obtain at least one set of configuration parameters for the shooting device; A first interface is displayed, which includes at least one set of configuration parameters and at least one option control, wherein the at least one option control is used to represent the selection of configuration parameters; In response to an operation on the target option control, the target configuration parameter corresponding to the target option control is determined to be the configuration parameter of the shooting device.
6. The method according to any one of claims 1 to 3, characterized in that, Before capturing an image based on the captured parameters, after analyzing and reasoning about the key elements based on a processing model to obtain the captured parameters, the method further includes: The second interface is displayed, which includes a confirmation control used to indicate the confirmation of applying the shooting parameters. In response to the operation of the confirmation control, an operation to capture an image based on the shooting parameters is triggered.
7. The method according to claim 6, characterized in that, The method further includes: The second interface is displayed, which includes a second preview image generated based on the shooting parameters.
8. A shooting device, characterized in that, include: The acquisition module is used to acquire the first preview image; The processing module is used to perform image recognition on the first preview image to obtain key elements of the first preview image, wherein the key elements are used to characterize the image information of the first preview image; The key elements are analyzed and reasoned based on the processing model to obtain the shooting parameters; The shooting module is used to capture images based on the shooting parameters.
9. A computer device, characterized in that, The computer device includes: One or more processors; The memory; and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the method of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to perform the steps of the method according to any one of claims 1 to 7.