Scalable and autonomous camera calibration system

Through a two-stage camera debugging process using multimodal LLM and RAG systems, image quality problems are automatically identified and resolved, and customized camera configuration solutions are generated. This solves the problems of time-consuming, labor-intensive, and poorly adaptable existing camera debugging processes, and achieves rapid and automated camera debugging and image quality improvement.

CN122349049APending Publication Date: 2026-07-07INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTEL CORP
Filing Date
2025-12-04
Publication Date
2026-07-07

AI Technical Summary

Technical Problem

The existing camera debugging process is time-consuming and labor-intensive, making it difficult to quickly adapt to the aesthetic preferences and image quality standards of different OEMs, and it is also prone to human error and deviation.

Method used

Employing a multimodal large language model (LLM) and retrieval-enhanced generation (RAG) system, a two-stage camera debugging process is implemented. The first stage identifies image quality issues, and the second stage autonomously generates a customized camera configuration solution.

Benefits of technology

It enables rapid and automated camera setup, reduces manual intervention, improves image quality consistency and adaptability, lowers costs, and shortens camera time to market.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122349049A_ABST
    Figure CN122349049A_ABST
Patent Text Reader

Abstract

The camera debugging process is a time-consuming and labor-intensive process. To address this problem, a camera debugging system including a multimodal large language model and a retrieval-augmented generation system can be implemented to intelligently and efficiently handle camera debugging tasks in real-time. The multimodal large language model can evaluate image quality and can be fine-tuned using high-quality labeled data and synthetically generated labeled data. The retrieval-augmented generation system can incorporate camera configuration knowledge into a vector database and can leverage the retrieved context to generate configuration solutions that address image quality issues identified by the multimodal large language model. The resulting camera debugging system is a unified process that can identify image quality issues and provide configuration solutions that address technical and aesthetic image quality issues.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] A camera is an optical system that captures and records light to create an image. A camera may include components such as a lens, a sensor, and a processing unit that processes the signals captured by the sensor. Cameras often face image quality issues or artifacts that impact user experience and product value. Attached Figure Description

[0002] The embodiments will be readily understood from the following detailed description taken in conjunction with the accompanying drawings. For ease of description, similar reference numerals denote similar structural elements. In the accompanying figures, embodiments are shown by way of example rather than limitation.

[0003] Figure 1 The laborious and time-consuming camera debugging process according to some embodiments of this disclosure is illustrated.

[0004] Figure 2 A camera debugging system involving a two-stage process is shown according to some embodiments of the present disclosure.

[0005] Figure 3 An analysis phase of a two-stage process is illustrated according to some embodiments of the present disclosure.

[0006] Figure 4 The solution generation phase of a two-stage process is illustrated according to some embodiments of the present disclosure.

[0007] Figure 5 This is a flowchart illustrating a method for debugging a camera according to some embodiments of the present disclosure.

[0008] Figure 6 An exemplary implementation of a single-modal large language model according to some embodiments of the present disclosure is shown.

[0009] Figure 7 An exemplary implementation of a multimodal large language model according to some embodiments of the present disclosure is shown.

[0010] Figure 8 This is a block diagram of an exemplary computing device according to some embodiments of the present disclosure. Detailed Implementation

[0011] Overview A camera system may include a lens, a sensor, and an image processing unit that processes the signals captured by the sensor. The sensor can capture raw images, and the image processing unit can receive the raw images and generate processed images, such as YUV or RGB images. The camera system can be tuned to mitigate image quality problems or reduce image quality artifacts in the processed images. Image quality problems may include inaccurate color reproduction leading to unnatural skin tones, limited dynamic range affecting shadow or highlight details, excessive noise in low light, and incorrect focus leading to blur. Image quality artifacts may include over-sharpening, incorrect white balance, saturation cropping, noisy images, incorrect lens shading correction, motion blur, etc.

[0012] Camera tuning may involve identifying and changing the available settings of a camera system. Modern cameras can be complex and may include many different types of settings. Available camera settings may include physical parameters (e.g., exposure time, aperture, focal length, etc.), sensor configurations (e.g., gain, pixel clock, power saving mode, etc.), and image processing and / or post-processing algorithm configurations (e.g., noise reduction parameters, sharpening parameters, tone mapping parameters, gamma correction parameters, color balance parameters). These settings can be controlled through hardware registers, camera software development kits, and / or user interfaces, where their configuration significantly affects image quality, noise levels, dynamic range, and color accuracy.

[0013] Camera tuning can be used to address the subjective aesthetic preferences (or goals) and / or various image quality standards of the original equipment manufacturer (OEM) of the camera system. Examples of subjective aesthetic preferences may include specific skin tone reproduction, with some preferring vibrant and reddish tones while others opt for more subtle yellow hues.

[0014] The diversity and complexity of image quality issues and objectives often lead image quality engineers and image processing engineers to design custom camera configuration solutions to address these issues and meet various goals. Therefore, there is no universal, "one-size-fits-all" camera configuration solution.

[0015] The camera setup process used to solve image quality problems and meet OEM preferences / goals can be a manual and laborious process that consumes a lot of time and resources. Figure 1This illustrates why camera setup can be a resource-intensive, time-consuming, and labor-intensive process. OEM 122 may have image quality standards and / or subjective aesthetic preferences 120. In 102, OEM 122 may use camera 128 to capture a test scene 126. In 104, OEM 122 obtains and evaluates the captured image 124. In 106, OEM 122 may report issues 132 (e.g., one or more image quality issues) to image quality engineer 182 while evaluating the captured image 124. In 108, image quality engineer 182 may attempt to replicate issues 132 in a laboratory environment and use camera 138 to capture a test scene 140 simulating test scene 126. In 110, image quality engineer 182 obtains and evaluates the captured image 134 of test scene 140. Replicating issues 132 reported by OEM 122 can be a complex task due to varying conditions and subjects in the test scene. In step 112, image quality engineer 182 can fine-tune camera settings 130 used by camera 138 (e.g., white balance, exposure, color profile, algorithm configuration, etc.) to capture test scene 140. In step 114, a captured image 136 of test scene 140 is obtained. Image quality engineer 182 can evaluate the captured image 136 to determine if camera settings 130 resolves problem 132. The process of fine-tuning camera settings 130 (including steps 112 and 114) can be repeated for test scene 140 until image quality engineer 182 determines that camera settings 130 resolves problem 132. Furthermore, image quality engineer 182 can repeat the fine-tuning process across a wide range of other test scenes besides test scene 140 to ensure that camera settings 130 do not adversely affect image quality in other settings. Once camera settings 130 have been verified and tested, in step 116, camera settings 130 are sent back to OEM 122 for approval.

[0016] Figure 1 The camera commissioning process illustrated can take up to 10 weeks and consume a significant workload for image quality engineers. The camera commissioning process can be particularly challenging for 3A problems (e.g., auto exposure, auto focus, and auto white balance) due to the dynamic nature of the algorithms controlling the image capture process and the feedback loops they create. Furthermore, the camera commissioning process can be challenging due to the complexity and difficulty of replicating the exact conditions under which image quality problems are initially observed by the OEM. The camera commissioning process is manual and can be expensive. This process cannot be quickly and easily scaled and deployed across many types of camera systems to provide customized camera configuration solutions.

[0017] To address this issue, a camera debugging system incorporating a multimodal large language model (LLM) and a retrieval-enhanced generation (RAG) system can be implemented to intelligently and efficiently handle camera debugging tasks in real time. The camera debugging system implements a two-stage camera debugging process. The first stage is the analysis stage. The second stage is the solution generation stage. The resulting camera debugging system is a unified process that can identify image quality problems and provide configurable solutions to address both technical and aesthetic image quality issues.

[0018] In the first phase, a multimodal LLM is implemented to autonomously evaluate image quality, compare images, identify artifacts, and provide insights into various image distortions. Engineering prompts for the task, including images and identifying image quality issues present in them, can be generated and fed into the multimodal LLM to obtain a response that includes the identified image quality issues. The multimodal LLM can be fine-tuned using high-quality labeled data, including synthetically generated labeled data. Synthetically generated labeled data can be produced with little or no human labor to enhance the fine-tuning of the multimodal LLM. Furthermore, applying multimodal LLM to image quality analysis tasks means that the task is less susceptible to human error and bias. This approach differs significantly from manual evaluation methods. Using a multimodal LLM allows for scalable and rapid responses to image quality issues, while addressing OEM feedback and adapting to their aesthetic preferences in real time.

[0019] In the second phase, the RAG system is implemented to autonomously generate customized camera configuration solutions. Having gained a deeper understanding of the image quality issues in the first phase, the LLM within the RAG system can autonomously generate camera configuration solutions that correct detected artifacts and / or align the camera output with a reference target image from the OEM. The RAG system can incorporate camera configuration knowledge into a vector database and leverage the retrieved context to generate configuration solutions that address the image quality issues identified by the multimodal LLM. Specifically, the RAG system can retrieve context using the image quality issues identified from the first phase. Based on the image quality issues identified in the response from the first phase, the query can be formatted and transformed into a query embedding. Using the query embedding, context can be retrieved from a vector database of embeddings containing camera configuration knowledge. Specifically, the vector database has embeddings generated for blocks of camera configuration knowledge. One or more embeddings closest to the query embedding can be retrieved, and one or more camera configuration knowledge blocks corresponding to one or more embeddings can be used as the context for retrieval. The retrieved context can include, for example, camera settings and / or parameters related to the identified image quality issues. The retrieved context can include camera configuration knowledge related to the query. Additional engineering hints can be generated, which may include the context of queries and retrievals, as well as additional tasks for determining camera configuration solutions. These additional engineering hints can be input into another LLM to obtain an additional response, which includes a specified camera configuration solution that addresses the image quality issues identified in the first stage. By incorporating camera configuration knowledge (e.g., camera processing unit specifications and / or manuals) into the RAG system, the additional LLM can autonomously generate customized camera configuration solutions that address the image quality issues using the retrieved context with relevant parameters and settings.

[0020] A two-stage camera tuning process bridges the gap between OEM goals and the resulting camera configuration solutions. In the first stage, a multimodal LLM is fine-tuned for expert-level image quality assessment using high-quality labeled data. In the second stage, an additional LLM is fed with relevant embeddings of camera configuration knowledge, allowing this additional LLM to generate a customized camera configuration solution. The result is a holistic camera tuning system designed to automatically identify and correct image quality issues, and to adapt in real-time to the subjective aesthetic preferences / goals of different OEMs. By integrating image analysis, real-time feedback, and continuous learning, the camera tuning system effectively achieves enhanced image quality tailored to specific needs without extensive manual intervention from specialized image quality engineers and image processing engineers. The holistic camera tuning system provides a unified solution for image quality assessment and camera reconfiguration, streamlining the camera tuning process.

[0021] In one aspect, the entire camera tuning system provides real-time OEM customization and image quality issue identification and resolution. Multimodal LLM effectively adapts to OEM aesthetic preferences and resolves image quality problems in real time. This approach bypasses the cumbersome and laborious process of replicating the OEM testing environment, such as... Figure 1 As shown in the diagram, a streamlined camera debugging process reduces the time and resources required to achieve optimal image quality. This process can be easily scaled to support diverse and changing OEM needs with minimal effort. The entire system provides real-time image quality troubleshooting for both OEMs and end-users of the camera system. The system's ability to operate automatically without human intervention allows it to provide real-time solutions within the camera system. Therefore, the system can be used in both offline and online camera debugging scenarios. In online camera debugging scenarios, the system proactively adjusts camera settings by instantly detecting and resolving image quality issues as the OEM or end-user captures images.

[0022] According to one aspect, by accessing and incorporating appropriate manuals, documentation, and knowledge data into the vector database, the RAG system in the second phase of the process can be extended to any camera. Therefore, the additional LLM within the RAG system can generate general, intelligent, and customized camera reconfiguration solutions for any camera, and can be easily updated by incorporating new camera configuration knowledge into the vector database.

[0023] According to one aspect, feedback from OEMs and / or end users on image quality issues can be collected and used as training data to fine-tune the multimodal LLM over time. This means that the multimodal LLM can continuously learn and improve based on new labeled data to refine and enhance its image quality assessment capabilities.

[0024] In one respect, the flexibility of multimodal LLM in handling different types of input allows camera debugging systems to adapt to a wide range and variations in preferences and reference targets, and to handle test images captured across different environmental conditions. The resulting system is scalable and adaptable.

[0025] Some advantages of improved camera tuning systems can include reduced burnout and workload for image quality engineers, faster camera system optimization, and scalable systems that can be adapted to or customized OEM-specific preferences / goals and various cameras. Camera tuning systems effectively combine the tasks of identifying image quality problems, generating customized solutions for specific cameras, and addressing technical and subjective aesthetic preferences / goals within a unified, autonomous system. In some cases, camera tuning systems may be less susceptible to human-related errors, oversights, biases, or mistakes when a trained machine learning model is identifying image quality problems and generating configuration solutions. Camera tuning systems can lead to consistently better camera image quality tuning results, faster time-to-market for camera systems, greater robustness to human error and bias, and potentially a more personalized camera experience for end users. Compared to… Figure 1 The reduction in manual labor and repeated testing in the camera setup process shown can lead to significant cost savings and improved satisfaction from OEMs and end users.

[0026] In one experiment, a camera tuning system was used to generate camera configuration solutions for sharpening boxes implemented in the camera system's processing unit. Using a query identifying an image quality problem (e.g., over-sharpening) within the RAG system, the RAG system had access to the camera system's documentation and tuning tools. The documentation and specifications (e.g., including parameter lists) were converted into embeddings in a vector database, and the RAG system retrieved ten (10) best matches as the retrieval context for the query. Hints were generated using the query and the retrieved context, and in response, the LLM in the RAG system responded by generating a detailed sequence of actions, precisely pinpointing the camera parameters to be adjusted to address the image quality problem. This experiment highlights the camera tuning system's ability to autonomously develop reconfiguration strategies, effectively resolving identified image quality problems without significant human intervention.

[0027] While many of the embodiments described herein are for camera commissioning, this disclosure contemplates that the teachings can be extended to commissioning other sensor systems, such as depth sensor systems, distance sensing systems, infrared sensing systems, etc.

[0028] Related work Figure 1 The diagram illustrates a less-than-ideal camera setup process involving manual tuning and calibration by image quality and image processing engineers, using standardized test charts, and software tools for image analysis. This process is labor-intensive, prone to human error, and difficult to adapt to the diverse aesthetic preferences of OEMs.

[0029] Some solutions use neural network deep learning methods to predict quality scores, but not LLM. These methods are limited by the scope of the dataset, the complexity of the real world, and struggle to adapt to different OEM aesthetic preferences. Furthermore, these methods do not address the image quality enhancement task.

[0030] In some solutions, off-the-shelf multimodal LLMs can provide some form of quality comparison, but lack a specialized and nuanced understanding of accurate image quality assessment for real-world tasks. Off-the-shelf multimodal LLMs typically fail to provide accurate image quality assessments. Furthermore, techniques such as few-lens learning and cue engineering have not consistently improved their performance in the field of image quality assessment. Moreover, off-the-shelf LLMs lack knowledge of specific camera configurations for particular cameras and therefore cannot generate camera configuration solutions to address image quality issues.

[0031] The various embodiments of scalable and autonomous camera tuning systems described herein differ from these solutions. Camera tuning systems provide a unified solution that can automatically adjust camera settings in real time and translate understanding of image quality into actionable profiles. Camera tuning systems can provide a comprehensive approach to image quality engineering tasks, thus offering a complete end-to-end system for improving image quality in camera systems.

[0032] Scalable and autonomous camera adjustment system The camera tuning system integrates LLM technology with advanced analytics, inference, and generation capabilities, such as multimodal LLM and RAG systems, to assess image quality issues and generate optimized camera configuration solutions. Figure 2 A camera debugging system involving a two-stage process is illustrated according to some embodiments of the present disclosure. The system operates in two stages (or phases): an analysis stage 280 and a solution generation stage 290.

[0033] Analysis phase 280 may include a prompt generator 206 and an analysis multimodal LLM 208. Figure 3 The diagram illustrates a detailed description of analysis phase 280. Solution generation phase 290 may include a solution generator RAG system 212. Figure 4 The diagram shows a detailed description of solution generation phase 290.

[0034] In step 202, the OEM 122 (or end user) can use camera 128 to capture test scene 126 and generate captured image 220. Captured image 220 can be provided as input to analysis stage 280 (in the form of visual input). One or more additional captured images can be provided as input to analysis stage 280.

[0035] As previously described, OEM 122 may have one or more subjective aesthetic preferences 120. In 204, OEM 122 may optionally provide subjective aesthetic preferences 120 as input to analysis phase 280 (in the form of text input). Subjective aesthetic preferences 120 may include one or more of the following: aesthetic preferences (or targets) and image quality standards. In some cases, subjective aesthetic preferences 120 may include one or more reference target images 282. In 204, OEM 122 may optionally provide one or more reference target images 282 as input to analysis phase 280 (in the form of visual input).

[0036] In some cases, OEM 122 may evaluate captured image 220 and provide one or more feedback comments about captured image 220. In 204, OEM 122 may optionally provide one or more feedback comments about captured image 220 as input (in the form of text input) to analysis phase 280. In some cases, OEM 122 may describe captured image 220 (e.g., regarding the environmental conditions under which captured image 220 was captured or one or more characteristics of one or more objects in captured image 220) and provide information about captured image 220. In 204, OEM 122 may optionally provide information about captured image 220 as input (in the form of text input) to analysis phase 280.

[0037] The cue generator 206 can receive one or more inputs from potentially different modalities and generate cue 224. Cue 224 can be input into the analytical multimodal LLM 208. The analytical multimodal LLM 208 can generate a response 222 in response to receiving and processing cue 224. Response 222 may include one or more identified image quality issues.

[0038] The multimodal LLM 208 can receive one or more captured images (e.g., captured image 220) as part of a cue 224, and is instructed in the cue 224 to evaluate the one or more captured images against a set of quality metrics and / or aesthetic goals. This evaluation can take into account image quality issues such as common image artifacts, color accuracy, exposure levels, and other technical parameters that contribute to the perceived quality of the image.

[0039] The multimodal LLM 208 analysis may, in some cases, receive OEM-specific aesthetic preferences (e.g., subjective aesthetic preference 120 and / or one or more reference target images 282) as part of cues 224, and take these preferences into account when evaluating one or more captured images. Subjective aesthetic preference 120 may be conveyed as text clearly stating the OEM's goals / purposes. One or more reference target images 282 visually convey the OEM's preferences and exemplify the desired image quality results.

[0040] In some cases, the analysis of the multimodal LLM 208 may receive feedback comments from the OEM 122 regarding one or more captured images as part of prompt 224, and these feedback comments may be taken into account when evaluating one or more captured images. The feedback comments may be conveyed as text clearly stating that the OEM 122 has specific image quality issues regarding one or more captured images.

[0041] Once the multimodal LLM 208 identifies one or more image quality issues (e.g., areas for improvement) in response 222, the camera debugging system moves to the solution generation phase 290. Solution generation phase 290 implements a solution generator RAG system 212 to intelligently propose camera configuration solutions that can address the identified image quality issues. The response 222 with one or more identified image quality issues can be formatted as a query to the solution generator RAG system 212. This query can be transformed into a query embedding by the solution generator RAG system 212.

[0042] The solution generator RAG system 212 has access to a comprehensive camera configuration knowledge base containing, for example, camera parameters, their impact on image quality, and documented solutions to common image quality problems. The camera configuration knowledge is chunked and transformed into embeddings, and stored in a vector database. The solution generator RAG system 212 can retrieve the context of a query from the vector database using query embeddings. The query (with one or more image quality problems) and the retrieved context can be combined to form a hint that can be used to prompt the LLM formulation in the solution generator RAG system 212 to produce a response with a camera configuration solution 230 that addresses the identified image quality problem and aligns with the subjective aesthetic preferences 120 of the OEM 122.

[0043] In 214, the solution generator RAG system 212 can provide camera configuration solutions 230 to OEM 122 or directly to camera 128 to reconfigure camera 128 in response 222 to address one or more identified image quality issues and potentially address OEM 122’s subjective aesthetic preferences 120.

[0044] In some cases, the analysis of the multimodal LLM 208 can generate a response 222 in the form of a query, which can be used in the solution generator RAG system 212. In such cases, the query may not need to be formatted by the solution generator RAG system 212.

[0045] In some cases, the solution generator RAG system 212 may receive one or more additional inputs besides the response 222 to help it retrieve context and / or generate hints that may lead to a better camera configuration solution 230. One or more additional inputs to the solution generator RAG system 212 may include one or more of the following: captured image 220, one or more feedback comments on the captured image 220, subjective aesthetic preferences 120, one or more reference target images 282, information about the camera 128, and information about the captured image 220.

[0046] Figure 3 An analysis phase of a two-stage process is illustrated according to some embodiments of this disclosure. A prompt generator 206 can generate a prompt 224. The prompt generator 206 may include one or more components: including a character 302, including a task 304, and including one or more images 386, which may include content or have content inserted into the prompt 224. The prompt generator 206 may utilize a prompt template including predefined content. The prompt 224 may be input into an analytical multimodal LLM 208 to generate a response 222. The prompt 224 may be input into a multimodal LLM such as the analytical multimodal LLM 208 to obtain a response 222 with one or more identified image quality issues.

[0047] The inclusion of one or more images 386 may include images such as the captured image 220 in prompt 224. In some cases, the inclusion of one or more images 386 may include one or more reference target images 282 in prompt 224.

[0048] Task 304 may be included in prompt 224. This task may specify or instruct the analysis of multimodal LLM 208 to determine image quality problems present in an image (e.g., captured image 220). In some embodiments, the task may include one or more types of possible image quality problems. In some embodiments, the task may include one or more characteristics of possible image quality problems. Examples of possible image quality problems may include lighting conditions, exposure, color balance, tone mapping, sharpness, and image noise. Where one or more reference target images 282 are included in prompt 224, the task may include instructions to compare the captured image with one or more reference target images 282. In some embodiments, the task may include instructions to compare multiple captured images and optionally one or more reference target images 282. In some embodiments, the task may include instructions to compare multiple captured images. In some embodiments, the task may include one or more of the following: aesthetic preferences and image quality criteria, such as preferences / goals associated with the OEM or end user of the camera system. In some embodiments, the task may include subjective aesthetic preferences 120.

[0049] Role 302 may include the role of Analytical Multimodal LLM 208, specifying that Analytical Multimodal LLM 208 is an Expert Image Quality Engineer.

[0050] Implementing the analytical multimodal LLM 208 is not straightforward. Various experiments using off-the-shelf multimodal LLMs have shown that they may struggle to produce responses that accurately identify image quality issues. This is not surprising, as off-the-shelf multimodal LLMs are trained to perform general image understanding tasks and may lack expert-level understanding of subtle image quality problems. To ensure that the analytical multimodal LLM 208 is robust and can accurately extract image quality issues, fine-tuning 314 and labeled data 312 are included to train and fine-tune the multimodal LLM used for the analytical multimodal LLM 208.

[0051] In some embodiments, fine-tuning 314 may involve updating the weights (or parameters) of the analytical multimodal LLM 208 using labeled data 312 to optimize the analytical multimodal LLM 208, thereby enabling the professional extraction of image quality issues.

[0052] One exemplary approach that can be implemented by fine-tuning 314 is Low-Rank Adaptation (LoRA). LoRA adds a small trainable rank factorization matrix to the existing weight matrix of the general multimodal LLM instead of modifying all the parameters of the general multimodal LLM. For example, in the attention layer of the transformer, LoRA can decompose a large weight matrix W into W + BA, where B and A are much smaller matrices. This reduces memory usage and training time while still allowing meaningful adaptation to the behavior of the general multimodal LLM. During training, only the LoRA parameters (matrices A and B) are updated, while the original weights of the multimodal LLM are frozen. During training, the multimodal LLM, along with the low-rank factorization matrix, processes batches of training samples from labeled data 312 and updates the LoRA parameters using gradient descent. For each example from labeled data 312, the multimodal LLM, along with the low-rank factorization matrix, makes a prediction using the current parameters, calculates the loss by comparing the prediction to the true label of the example, computes the loss gradient relative to the LoRA parameters using backpropagation, and updates the LoRA parameters in the direction that reduces the loss. The training process can be repeated on batches of samples from labeled data 312 until convergence to the optimal LoRA parameters.

[0053] The labeled data 312 may include a comprehensive and diverse fine-tuned dataset, which was developed to enhance the ability of the multimodal LLM 208 to extract image quality issues.

[0054] In some cases, the label data 312 includes synthetically created label data, which can be generated by synthetic label data creation 310. Synthetic label data creation 310 can artificially induce image artifacts and distortion by configuring the camera with known camera configurations / settings that are detrimental to image quality. Synthetic label data creation 310 can artificially induce image artifacts and distortion by distorting the captured image through post-processing (e.g., introducing incorrect white balance gain, adding lens shading, etc.). Images, along with one or more corresponding image quality issues as labels, can be generated by synthetic label data creation 310 and stored in the label data 312. In some cases, images without one or more causative image quality issues can be stored in the label data 312.

[0055] In some cases, labeled data 312 includes collected labeled data, which may be collected by labeled data collection 330. Label data collection 330 may collect historical image quality problem reports, such as reports of specific image quality problems in past captured images and detailed descriptions from OEMs, and store these historical image quality problem reports in labeled data 312. Label data collection 330 may collect professional evaluations from image quality engineers or image processing engineers for captured images and store the captured images and professional evaluations in labeled data 312. Including professional evaluations in labeled data 312 can leverage qualitative insights and annotations from experienced engineers as high-quality labeled data, thus providing high-quality training data in labeled data 312. Label data collection 330 may collect images captured using and / or without a camera configuration solution previously handcrafted or autonomously generated to address one or more image quality problems, and store the images and corresponding image quality problems in labeled data 312. Label data collection 330 may collect image evaluations from internal data sources and include these evaluations in labeled data 312. Including source image evaluations in labeled data 312 diversifies the training dataset within it. In one example, an internal data source provides 8600 paired data samples from an internal camera professional review database, and 450 data samples are selected and stored in labeled data 312. Collecting labeled data 330 can collect comparative images, including side-by-side image comparisons of cameras under different configurations / settings or the same camera, and annotate them with one or more image quality issues, storing the comparative images in labeled data 312. Collecting labeled data 330 can also collect an open-source image quality training dataset and store it in labeled data 312.

[0056] In one example, when a prompt with three images and a task for comparing the images is input into the Analysis Multimodal LLM 208, the Analysis Multimodal LLM 208 generates the following response: Figure 4 The solution generation stage of a two-stage process is illustrated according to some embodiments of the present disclosure. The solution generation stage integrates RAG technology. Specifically, Figure 4 An exemplary implementation of a solution generator RAG system 212 with RAG technology is shown.

[0057] RAG (Related Aspects of Graphs) technology enhances LLM (Limited Module Management) by combining a model with a knowledge retrieval system. Knowledge can first be chunked into smaller pieces (e.g., using techniques like sliding windows or semantic partitioning). These chunks can be transformed into dense vector embeddings using a transformer-based model or LLM. Embeddings can be stored in a vector database for efficient similarity search. The vector database uses a high-dimensional vector space to store embeddings, where each document or item is represented as a numeric vector. When performing a similarity search, the vector database can use a specialized indexing structure to avoid comparing the query embedding with every stored embedding. These indexes create a graph-like structure where similar vectors are concatenated, allowing the search to quickly traverse to the most relevant candidates. The database then sorts the closest matches using a distance metric such as cosine similarity and can return the top k embeddings.

[0058] For more specific reference Figure 4 The solution generator RAG system 212 can utilize rich camera-specific configuration knowledge (shown as configuration knowledge 402) to generate a camera configuration solution 230. Configuration knowledge 402 may include one or more of the following: parameter names, their corresponding value ranges, camera parameter functionality, detailed image processing algorithm specifications, methods used by image quality tools, previous image quality problems and corresponding solutions, user manuals, image quality tool code, camera register descriptions, camera system calibration tool specifications, reference target images, previously captured images and corresponding solutions, comparison images illustrating one or more image quality problems and the configuration settings leading to the problems, etc. In some embodiments, configuration knowledge 402 includes camera parameter names, camera parameter value ranges for the camera parameter names, camera functionality, image processing algorithm specifications, image quality tool methods, camera manuals, register configurations, and firmware configuration code.

[0059] Configuration knowledge 402 is encoded as embeddings through conversion to embeddings 404, and these embeddings can be stored in the vector database 406. Conversion to embeddings 404 can use a transformer-based model or LLM to convert blocks of configuration knowledge 402 into embeddings. An embedding is a retrieveable format that the solution generator RAG system 212 can access and apply to its configuration solution generation process. Specifically, multiple most prominent or relevant embeddings and corresponding fragments or blocks of knowledge / information can be retrieved from the vector database 406. These fragments or blocks of knowledge / information can be used as context, and the LLM 480 can use the retrieved context to generate accurate and effective camera configuration solutions.

[0060] Queries, such as questions, can be submitted to the RAG system to initiate the knowledge retrieval process and retrieve the most prominent or relevant embeddings from the vector database 406. The information fragments corresponding to the most prominent or relevant embeddings can be used as context. Figure 4 In the example, the query may be based on one or more identified image quality issues in response 222. Query formatting 410 may format the query based on one or more identified image quality issues in response 222. Query formatting 410 may reformulate the identified image quality issues in response 222 as questions. In some cases, one or more additional pieces of information may be appended to or added to the query through query formatting 410. One or more additional pieces of information may include one or more of aesthetic preferences, image quality standards, and subjective aesthetic preferences 120. One or more additional pieces of information may include one or more of the following: one or more reference target images 282 and captured images 220. One or more additional pieces of information may include one or more of the following: information about the camera and information about captured images 220.

[0061] Converting to Query Embedding 412 converts the query generated in Query Formatting 410 into a query embedding. Converting to Query Embedding 412 can utilize (the same or different) transformer-based models or LLMs used to generate embeddings in Vector Database 406. Converting to Query Embedding 412 transforms the query into the same embedding space or vector space as the embeddings in Vector Database 406.

[0062] Using query embeddings, retrieval context 492 can retrieve context from vector database 406. Retrieval context 492 can perform similarity searches using query embeddings. Similarity can be measured using cosine similarity or dot product. Retrieval context 492 can retrieve multiple embeddings of configuration knowledge 402 that best match the query embedding. Retrieval context 492 retrieves the top k most relevant blocks (e.g., those with the top k highest similarity scores) corresponding to the top k closest or most similar embeddings of the query embedding. The retrieval context may include the top k most relevant configuration knowledge blocks 402. Optionally, filtering and / or reordering can be performed by retrieval context 492 when generating context.

[0063] The prompt generator 494 can receive a query generated or produced by the query formatting 410 and a retrieval context obtained from the retrieval context 492, and generate an additional prompt 498 with the query, context, and additional task. This additional task can specify or instruct the LLM 480 to determine a configuration solution. The additional prompt 498 can be input into the LLM 480 to obtain an additional response with a specified configuration solution that resolves one or more identified image quality problems in response 222.

[0064] The prompt generator 494 may include one or more components: including a role 444, including a task 446, including a query 448, and including a context 450, which may include content or insert content into additional prompts 498. The prompt generator 494 may utilize prompt templates that include predefined content. Additional prompts 498 may be input into the LLM 480 to generate additional responses with the camera configuration solution 230.

[0065] Role 444 may include the role of LLM 480, specifying that LLM 480 is an expert image processing engineer.

[0066] Query 448 can include the query generated by query format 410 in additional prompt 498.

[0067] The context 450 may include the search context obtained from the search context 492 in an additional prompt 498.

[0068] Task 304 may include an additional task in a separate prompt 498. This additional task may specify or instruct the LLM 480 to determine a camera configuration solution. In some embodiments, the additional task includes instructions to determine a configuration solution for a specific box in the image processing pipeline (e.g., a filtering box, an image processing box, a post-processing box, etc.). In some embodiments, the additional task includes instructions to output the configuration solution in the form of one or more operations. In some embodiments, the additional task includes instructions to output the configuration solution in the form of one or more register values ​​and one or more register addresses to write one or more register values. In some embodiments, the additional task includes instructions to output the configuration solution in the form of one or more application programming interface (API) function calls or code having one or more API function calls.

[0069] Using additional clues 498 enriched by the retrieved context, the LLM 480 can synthesize camera configuration solutions that address one or more identified image quality issues in response 222. The retrieval context in additional clue 498 can base the responses generated by the LLM 480 on configuration knowledge 402, while maintaining the LLM 480's ability to reason and synthesize camera configuration solutions in the realm of expert image quality analysis. Furthermore, the vector database 406 can be updated with real-time knowledge updates and new camera configuration knowledge without requiring retraining or fine-tuning of the LLM 480.

[0070] In one example, the LLM 480 produces the following further response: LLM 480 can be a standalone large language model. LLM 480 can be a unimodal LLM. LLM 480 can be a multimodal LLM. LLM 480 can be the same model as the multimodal LLM 208 analysis model. LLM 480 can share parts of the multimodal LLM 208 analysis model. LLM 480 can be... Figure 3 Fine-tuning 314 Figure 3 The 312 labeled data is used for fine-tuning.

[0071] In this article, prompts used as inputs to LLM (such as...) Figure 2-3 Hint 224 Figure 4 Additional note 498) may include content with one or more modalities.

[0072] Exemplary methods for camera debugging Figure 5 This is a flowchart illustrating a method for debugging a camera according to some embodiments of the present disclosure. Method 500 can be... Figure 2-4 Method 500 can be executed using one or more components as shown. Figure 8 It is executed by a computing device such as the 800 computing device.

[0073] In error 502, a prompt is generated. This prompt may include an image and a task to identify image quality issues present in the image.

[0074] In the 504 error, the prompt is fed into the multimodal LLM to obtain a response that includes the identified image quality issues.

[0075] In 506, the query is formatted based on the image quality issues identified in the response.

[0076] In 508, the query is transformed into a query embedding.

[0077] In 510, query embeddings are used to retrieve context from a vector database of embeddings containing camera configuration knowledge. This context contains one or more pieces of camera configuration knowledge relevant to the query.

[0078] In 512, additional suggestions are generated. These additional suggestions have additional tasks such as querying, providing context, and determining the configuration solution.

[0079] In 514, additional prompts are entered into another LLM to obtain an additional response, which includes a specified configuration solution to resolve the image quality issues identified.

[0080] In some embodiments, alternative operations may be performed, including: generating a prompt for a task having an image and identifying an image quality problem present in the image; inputting the prompt into a multimodal large language model to obtain a response including the identified image quality problem; using the identified image quality problem, retrieving context from an embedded vector database having camera configuration knowledge, the context having one or more pieces of camera configuration knowledge related to the identified image quality problem; generating an additional prompt having the identified image quality problem, the context, and an additional task for determining a configuration solution; and inputting the additional prompt into another large language model to obtain an additional response including a specified configuration solution for resolving the identified image quality problem.

[0081] In some embodiments, alternative operations may be performed, including: generating a prompt for a task having an image and identifying an image quality problem present in the image; inputting the prompt into a multimodal large language model to obtain a response including the identified image quality problem; formatting a query based on the identified image quality problem in the response; retrieving context from a vector database having camera configuration knowledge, the context having one or more pieces of camera configuration knowledge related to the query; generating additional prompts for the query, the context, and an additional task for determining a configuration solution; and inputting the additional prompts into another large language model to obtain an additional response including a specified configuration solution for resolving the identified image quality problem.

[0082] Large language model The various sections of the description discuss the use of LLM. Regarding... Figure 6-7 The description illustrates some implementations of LLM, as well as details about transformer-based models.

[0083] Figure 6 An exemplary implementation of a single-modal LLM 600 according to some embodiments of the present disclosure is shown. The single-modal LLM 600 can receive input 680 and generate output 682. Input 680 can be single-modal, such as image / vision, text / text, audio, sensor signal, etc. Input 680 can include an input sequence or a series of inputs. The single-modal LLM 600 can process input 680 into digital tokens through tokenization (not explicitly shown), transformation of input 680 (e.g., words and subwords of text input, batch image input, cropping of audio input, portions of an input sequence, etc.), and an input encoder 666 can convert the digital tokens into a continuous vector representation, shown as input embedding 668. In some cases, the token order can be captured by the input encoder 666 including positional encoding. Input embedding 668 can include semantic information about each token in the input sequence and its context.

[0084] Input embedding 668 can be fed into transformer 670. Transformer 670 can generate output 682 token-by-token, where each new token is conditioned on the input sequence and previously generated tokens. Transformer 670 may include an internal encoder 602 and an internal decoder 612. Each token can flow through the entire encoder-decoder pipeline of transformer 670 before generating the next token.

[0085] The internal encoder 602 and internal decoder 612 include corresponding transformer layers. The transformer layers may include one or more of the following: self-attention mechanisms, cross-attention mechanisms, masked self-attention mechanisms, feedforward neural networks, multi-head self-attention mechanisms, etc. In some implementations, the transformer layer may include two sub-layers: multi-head attention and a positional feedforward network, each wrapped with residual connections and layer normalization. Multi-head attention splits the input into multiple heads, which independently perform scaled dot-product attention using queries, keys, and values ​​before recombining the results. The feedforward network may include two linear transformations with an activation function between them, independently processing features at each location. Layer normalization helps stabilize training by normalizing features across embedding dimensions, while residual connections allow direct gradient flow and help preserve information through deep networks.

[0086] Inside transformer 670, internal encoder 602 can use self-attention mechanisms in one or more transformer layers 604 to process process input embeddings 668 to construct contextual representations, where each token pays attention to all other input tokens. These encoded representations are then passed as state 644 to internal decoder 612.

[0087] The internal encoder 602 may include one or more transformer layers 604. In some implementations, the transformer layers 604 in the internal encoder 602 may include a multi-head self-attention mechanism followed by a feedforward neural network. In self-attention, the representation of each label is updated by paying attention to all other labels in the input sequence, where the attention weights are determined by the projection of learned key query values. Multiple attention heads capture different types of relationships between labels. The feedforward network then processes these attention-weighted representations through two linear transformations with non-linear activations in between, allowing the model to transform the information collected through attention.

[0088] Inside transformer 670, internal decoder 612 may use cross-attention mechanisms in one or more transformer layers 614 to align each output token with the relevant portion of the encoded input (passed from internal encoder 602 to internal decoder 612 as state 644). Internal decoder 612 may include masked self-attention mechanisms in one or more transformer layers 614 to prevent looking ahead to future tokens during generation.

[0089] The internal decoder 612 may include one or more transformer layers 614. In some implementations, the transformer layer 614 may include three components: masked self-attention, cross-attention, and a feedforward neural network. Masked self-attention operates similarly to encoder self-attention, but prevents tokens from focusing on future positions during generation. Cross-attention allows decoder tokens to focus on encoder output, where decoder queries focus on encoder keys and values. The cross-attention mechanism bridges the input and output sequences, allowing each generated token to extract information from the associated input tokens. The feedforward neural network then processes these combined representations.

[0090] The encoder-decoder interaction in the transformer 670 allows the model to learn complex relationships between input and output sequences while maintaining the appropriate generation order.

[0091] Figure 7 Exemplary implementations of a multimodal LLM 700 according to some embodiments of the present disclosure are shown. The multimodal LLM 700 may differ from the LLM 600 because the multimodal LLM 700 receives multiple inputs with different modalities. As illustrated, the multimodal LLM 700 may receive inputs 780 and 790. For example, input 780 may include text input, and input 790 may include image input. It is envisioned that the multimodal LLM 700 may receive inputs having combinations of the modalities described herein. Corresponding input encoders may be provided to process the respective inputs to generate corresponding input embeddings. Input encoder 710 may process input 790 to generate input embedding 792. Input encoder 720 may process input 780 to generate input embedding 782. For each modality, the input encoders (e.g., input encoder 710 and input encoder 720) may be dedicated or implemented differently. For text input, the input encoder may use tokenization and embedding layers to generate the input embedding. For image / visual input, a visual transformer or convolutional neural network can be implemented in the input encoder to produce the input embedding. For audio input, a specialized audio processing architecture can be implemented in the input encoder to produce the input embedding. The corresponding input encoder can transform the (raw) input into a common embedding space.

[0092] In some embodiments, the multimodal LLM 700 includes a shared backbone network 730 to achieve modality fusion and generate joint input embeddings 732. The shared backbone network 730 may include one or more transformer layers. In some implementations, inputs from different modalities are projected by the shared backbone network 730 to have compatible dimensions and semantic structures. The shared backbone network 730 may use a cross-attention mechanism that processes input embeddings regardless of their original form. The shared backbone network 730 may include cross-attention transformer layers that can extract cross-modal relationships. In one example, the shared backbone network 730 may associate visual features with corresponding textual descriptions while maintaining modality-specific information. Additional projection layers and normalization techniques in the shared backbone network 730 can help manage the different statistical properties of embeddings from each modality, ensuring balanced contributions from all input types during joint processing.

[0093] The combined input 732 can be provided as input to the converter 670 to ultimately produce the output 722.

[0094] This disclosure envisions other architectures that can be implemented to provide modality fusion. For example, the shared backbone network 730 can be omitted, and the input embeddings 782 and 792 can be attached together and used as inputs to the transformer 670 (and the internal encoder 602 can play the role of modality fusion).

[0095] Exemplary computing device Figure 8 These are block diagrams of devices or systems according to some embodiments of the present disclosure, such as an exemplary computing device 800. One or more computing devices 800 may be used to implement the functionality described herein in conjunction with the accompanying drawings. Figure 8 The diagram illustrates that multiple components may be included in the computing device 800, but any one or more of these components may be omitted or duplicated to suit the application. In some embodiments, some or all of the components included in the computing device 800 may be attached to one or more motherboards. In some embodiments, some or all of these components are manufactured onto a single system-on-a-chip (SoC) die. Additionally, in various embodiments, the computing device 800 may not include... Figure 8The computing device 800 may include one or more components as shown, and may include interface circuit modules for coupling to one or more components. For example, the computing device 800 may not include the display device 806, but may include display device interface circuit modules (e.g., connector and driver circuit modules) to which the display device 806 may be coupled. In another set of examples, the computing device 800 may not include the audio input device 818 or the audio output device 808, but may include audio input or output device interface circuit modules (e.g., connector and support circuit modules) to which the audio input device 818 or the audio output device 808 may be coupled.

[0096] Computing device 800 may include processing device 802 (e.g., one or more processing devices, one or more processing devices of the same type, or one or more processing devices of different types). Processing device 802 may include electronic circuit modules that process electronic data from data storage elements (e.g., registers, memories, resistors, capacitors, qubit units) to convert the electronic data into other electronic data that can be stored in registers and / or memories. Examples of processing device 802 may include CPUs, GPUs, quantum processors, machine learning processors, artificial intelligence processors, neural network processors, artificial intelligence accelerators, application-specific integrated circuits (ASICs), analog signal processors, analog computers, microprocessors, digital signal processors, field-programmable gate arrays (FPGAs), tensor processing units (TPUs), data processing units (DPUs), etc.

[0097] Computing device 800 may include memory 804, which itself may include one or more memory devices, such as volatile memory (e.g., DRAM), non-volatile memory (e.g., read-only memory (ROM)), high-bandwidth memory (HBM), flash memory, solid-state memory, and / or hard disk drive. Memory 804 includes one or more non-transitory computer-readable storage media. In some embodiments, memory 804 may include memory sharing a die with processing device 802.

[0098] In some embodiments, memory 804 includes one or more non-transitory computer-readable media storing instructions executable to perform... Figure 2-7 And the operations described in this article, for example Figure 5 Method 500 is shown.

[0099] Memory 804 may store instructions encoding one or more exemplary portions. Exemplary portions (such as one or more portions of analysis phase 280 and one or more portions of solution generation phase 290) may be encoded as instructions and stored in memory 804. Exemplary portions may include one or more components of: synthetic tag data creation 310, fine-tuning 314, tag data collection 330, cue generator 206, analytical multimodal LLM 208, and solution generator RAG system 212. Instructions stored in one or more non-transitory computer-readable media may be executed by processing device 802.

[0100] In some embodiments, memory 804 may store data, such as data structures, binary data, bits, metadata, files, blobs, etc., as shown in the accompanying drawings and described herein. Exemplary data described herein (e.g., hints, queries, context, embeddings, vector databases, responses, tagged data, images, text, etc.) may be stored in memory 804.

[0101] In some embodiments, memory 804 may store one or more machine learning models (and / or portions thereof) used in the LLM and encoder described herein. Memory 804 may store training data used to train one or more machine learning models. In one example, memory 804 may store... Figure 3 The labeled data 312 is shown. Memory 804 can store input data, output data, intermediate outputs, and intermediate inputs of one or more machine learning models. Memory 804 can store instructions for performing one or more operations of the machine learning model. Memory 804 can store one or more parameters used by the machine learning model. Memory 804 can store information encoding how the processing units of the machine learning model are interconnected.

[0102] In some embodiments, computing device 800 may include communication device 812 (e.g., one or more communication devices). For example, communication device 812 may be configured to manage wired and / or wireless communications to transmit data to and from computing device 800. The term “wireless” and its derivatives may be used to describe circuits, apparatuses, systems, methods, technologies, communication channels, etc., which can transmit data using modulated electromagnetic radiation over a non-solid medium. This term does not imply that the associated apparatus does not contain any wiring, although in some embodiments they may not contain any wiring. Communication device 812 may implement any of a variety of wireless standards or protocols, including but not limited to Institute of Electrical and Electronics Engineers (IEEE) standards, including Wi-Fi (IEEE 802.10 series), IEEE 802.16 standards (e.g., IEEE 802.16-2005 amendments), Long Term Evolution (LTE) projects, and any amendments, updates, and / or revisions (e.g., Advanced LTE projects, Ultra Mobile Broadband (UMB) projects (also known as “3GPP2”), etc.). IEEE 802.16 compliant Broadband Wireless Access (BWA) networks are commonly referred to as WiMAX networks, an acronym for Global Microwave Access, and are certification marks for products that have passed conformance and interoperability testing according to the IEEE 802.16 standard. Communication device 812 can operate according to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Universal Mobile Telecommunications System (UMTS), High-Speed ​​Packet Access (HSPA), Evolved HSPA (E-HSPA), or LTE networks. Communication device 812 can operate according to Enhanced Data GSM Evolution (EDGE), GSM EDGE Radio Access Network (GERAN), Universal Terrestrial Radio Access Network (UTRAN), or Evolved UTRAN (E-UTRAN). Communication device 812 can operate according to Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Digital Enhanced Cordless Telecommunications (DECT), Evolved Data Optimization (EV-DO) and its derivative protocols, as well as any other wireless protocols designated as 3G, 4G, 5G, and above. In other embodiments, the communication device 812 may operate according to other wireless protocols. The computing device 800 may include an antenna 822 to facilitate wireless communication and / or receiving other wireless communications (e.g., radio frequency transmissions). The computing device 800 may include receiver circuitry and / or transmitter circuitry. In some embodiments, the communication device 812 may manage wired communication, such as electrical, optical, or any other suitable communication protocol (e.g., Ethernet). As described above, the communication device 812 may include multiple communication chips.For example, the first communication device 812 may be dedicated to short-range wireless communication such as Wi-Fi or Bluetooth, while the second communication device 812 may be dedicated to long-range wireless communication such as Global Positioning System (GPS), EDGE, GPRS, CDMA, WiMAX, LTE, EV-DO, etc. In some embodiments, the first communication device 812 may be dedicated to wireless communication, while the second communication device 812 may be dedicated to wired communication.

[0103] The computing device 800 may include a power source / power circuit module 814. The power source / power circuit module 814 may include one or more energy storage devices (e.g., batteries or capacitors) and / or circuit modules for coupling components of the computing device 800 to an energy source (e.g., DC power, AC power, etc.) separate from the computing device 800.

[0104] The computing device 800 may include a display device 806 (or a corresponding interface circuit module, as described above). The display device 806 may include any visual indicator, such as a head-up display, a computer monitor, a projector, a touch screen display, a liquid crystal display (LCD), a light-emitting diode display, or a flat panel display.

[0105] The computing device 800 may include an audio output device 808 (or a corresponding interface circuit module, as described above). The audio output device 808 may include any device that generates auditory indicators, such as a speaker, headphones, or earphones.

[0106] The computing device 800 may include an audio input device 818 (or a corresponding interface circuit module, as described above). The audio input device 818 may include any device that generates a signal representing sound, such as a microphone, microphone array, or digital musical instrument (e.g., a musical instrument with a Musical Instrument Digital Interface (MIDI) output).

[0107] The computing device 800 may include a GPS device 816 (or a corresponding interface circuit module, as described above). As is known in the art, the GPS device 816 can communicate with a satellite-based system and can receive the location of the computing device 800.

[0108] The computing device 800 may include a sensor 830 (or one or more sensors). The computing device 800 may include corresponding interface circuitry, as described above. The sensor 830 can sense physical phenomena and convert them into electrical signals that can be processed by, for example, the processing device 802. Examples of the sensor 830 may include: capacitive sensors, inductive sensors, resistive sensors, electromagnetic field sensors, light sensors, cameras, imagers, microphones, pressure sensors, temperature sensors, vibration sensors, accelerometers, gyroscopes, strain sensors, humidity sensors, distance sensors, time-of-flight sensors, pH sensors, particle sensors, air quality sensors, chemical sensors, gas sensors, biosensors, ultrasonic sensors, scanners, etc.

[0109] The computing device 800 may include another output device 810 (or a corresponding interface circuit module, as described above). Examples of other output devices 810 may include audio codecs, video codecs, printers, wired or wireless transmitters for providing information to other devices, haptic output devices, gas output devices, vibration output devices, lighting output devices, home automation controllers, or additional storage devices.

[0110] The computing device 800 may include another input device 820 (or a corresponding interface circuit module, as described above). Examples of the other input device 820 may include an accelerometer, gyroscope, compass, image capture device, keyboard, cursor control device such as mouse, stylus, touchpad, barcode reader, quick response (QR) code reader, any sensor, or radio frequency identification (RFID) reader.

[0111] The computing device 800 can have any desired form factor, such as a handheld or mobile computer system (e.g., a cellular phone, smartphone, mobile internet device, music player, tablet computer, laptop computer, netbook computer, personal digital assistant (PDA), personal computer, remote control, wearable device, helmet, glasses, footwear, electronic clothing, etc.), desktop computer system, server or other networked computing component, printer, scanner, monitor, set-top box, entertainment control unit, vehicle control unit, digital camera, digital video recorder, Internet of Things device, or wearable computer system. In some embodiments, the computing device 800 can be any other electronic device that processes data.

[0112] Select Example Example 1 provides a method comprising: generating a prompt having an image and a task for identifying an image quality problem present in the image; inputting the prompt into a multimodal large language model to obtain a response including the identified image quality problem; formatting a query based on the identified image quality problem in the response; converting the query into a query embedding; retrieving context from a vector database of embeddings having camera configuration knowledge, the context having one or more pieces of camera configuration knowledge related to the query; generating additional prompts having the query, the context, and an additional task for determining a configuration solution; and inputting the additional prompts into another large language model to obtain an additional response including a specified configuration solution for resolving the identified image quality problem.

[0113] Example 2 provides the method described in Example 1, wherein the hint also includes the role of a multimodal large language model, specifying that the multimodal large language model is an expert image quality engineer.

[0114] Example 3 provides the method described in Example 1 or 2, wherein the task includes one or more possible image quality problems.

[0115] Example 4 provides a method described in any of Examples 1-3, wherein the task includes one or more features of possible image quality problems.

[0116] Example 5 provides the method described in Example 3 or 4, wherein possible image quality issues include lighting conditions, exposure, color balance, tone mapping, sharpness, and image noise.

[0117] Example 6 provides a method described in any of Examples 1-5, wherein the task includes instructions to compare an image with a reference target image.

[0118] Example 7 provides a method described in any of Examples 1-6, wherein the task includes one or more of aesthetic preferences and image quality standards.

[0119] Example 8 provides a method described in any of Examples 1-7, wherein formatting the query based on the identified image quality problem includes restating the identified image quality problem as a question.

[0120] Example 9 provides a method for any of Examples 1-8, wherein formatting the query based on the identified image quality problem includes appending one or more of aesthetic preferences and image quality criteria to the identified image quality problem.

[0121] Example 10 provides a method described in any of Examples 1-9, wherein retrieving context from a vector database includes retrieving multiple embeddings of camera configuration knowledge that most closely match the query embedding.

[0122] Example 11 provides the method described in any one of Examples 1-10, wherein the camera configuration knowledge includes one or more of the following: camera parameter names, camera parameter value ranges for camera parameter names, camera functionality, image processing algorithm specifications, image quality tool methods, camera manuals, register configuration, and firmware configuration code.

[0123] Example 12 provides the method described in any of Examples 1-11, wherein the additional hint includes the role of a multimodal large language model, specifying that the multimodal large language model is an expert image processing engineer.

[0124] Example 13 provides the method described in any of Examples 1-12, wherein an additional task includes instructions for determining a configuration solution for a box in an image processing pipeline.

[0125] Example 14 provides a method described in any of Examples 1-13, wherein an additional task includes outputting instructions for configuring the solution in the form of one or more operations.

[0126] Example 15 provides the method described in any of Examples 1-14, wherein an additional task includes outputting instructions to write one or more register values ​​in the form of one or more register values ​​and one or more register addresses.

[0127] Example 16 provides the method described in any of Examples 1-15, wherein an additional task includes outputting instructions for configuring the solution in the form of one or more application programming interface function calls.

[0128] Example 17 provides one or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause one or more processors to: generate a prompt having an image and a task identifying an image quality problem present in the image; input the prompt into a multimodal large language model to obtain a response including the identified image quality problem; format a query based on the identified image quality problem in the response; convert the query into a query embedding; retrieve context from a vector database of embeddings having camera configuration knowledge, the context having one or more pieces of camera configuration knowledge related to the query; generate additional prompts having the query, the context, and an additional task for determining a configuration solution; and input the additional prompts into another large language model to obtain an additional response including a specified configuration solution for resolving the identified image quality problem.

[0129] Example 18 provides one or more non-transitory computer-readable media as described in Example 17, wherein the hint also includes the role of a multimodal large language model, specifying that the multimodal large language model is an expert image quality engineer.

[0130] Example 19 provides one or more non-transitory computer-readable media as described in Example 17 or 18, wherein the task includes one or more types of possible image quality problems.

[0131] Example 20 provides one or more non-transitory computer-readable media as described in any of Examples 17-19, wherein the task includes one or more features of possible image quality problems.

[0132] Example 21 provides one or more non-transitory computer-readable media as described in Example 19 or 20, wherein possible image quality problems include lighting conditions, exposure, color balance, tone mapping, sharpness, and image noise.

[0133] Example 22 provides one or more non-transitory computer-readable media as described in any of Examples 17-21, wherein the task includes instructions for comparing an image and a reference target image.

[0134] Example 23 provides one or more non-transitory computer-readable media as described in any of Examples 17-22, wherein the task includes one or more of aesthetic preferences and image quality standards.

[0135] Example 24 provides one or more non-transitory computer-readable media as described in any of Examples 17-23, wherein formatting a query based on an identified image quality problem includes restating the identified image quality problem as a question.

[0136] Example 25 provides one or more non-transitory computer-readable media as described in any of Examples 17-24, wherein formatting a query based on an identified image quality problem includes attaching one or more of aesthetic preferences and image quality criteria to the identified image quality problem.

[0137] Example 26 provides one or more non-transitory computer-readable media as described in any of Examples 17-25, wherein retrieving context from a vector database includes retrieving multiple embeddings of camera configuration knowledge that most closely matches the query embedding.

[0138] Example 27 provides one or more non-transitory computer-readable media as described in any of Examples 17-26, wherein the camera configuration knowledge includes one or more of the following: camera parameter names, camera parameter value ranges for camera parameter names, camera functionality, image processing algorithm specifications, image quality tool methods, camera manuals, register configurations, and firmware configuration code.

[0139] Example 28 provides one or more non-transitory computer-readable media as described in any of Examples 17-27, wherein the additional hint includes the role of a multimodal large language model, specifying that the multimodal large language model is an expert image processing engineer.

[0140] Example 29 provides one or more non-transitory computer-readable media as described in any of Examples 17-28, wherein an additional task includes instructions for determining a configuration solution for boxes in an image processing pipeline.

[0141] Example 30 provides one or more non-transitory computer-readable media as described in any of Examples 17-29, wherein an additional task includes outputting instructions for configuring the solution in the form of one or more operations.

[0142] Example 31 provides one or more non-transitory computer-readable media as described in any of Examples 17-30, wherein an additional task includes outputting instructions for writing one or more register values ​​in the form of one or more register values ​​and one or more register addresses.

[0143] Example 32 provides one or more non-transitory computer-readable media as described in any of Examples 17-31, wherein an additional task includes outputting instructions for configuring the solution in the form of one or more application programming interface function calls.

[0144] Example 33 provides an apparatus including one or more processors; and one or more non-transitory computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to: generate a prompt having an image and a task identifying an image quality problem present in the image; input the prompt into a multimodal large language model to obtain a response including the identified image quality problem; format a query based on the identified image quality problem in the response; convert the query into a query embedding; retrieve context from a vector database of embeddings having camera configuration knowledge, the context having one or more pieces of camera configuration knowledge related to the query; generate additional prompts having the query, the context, and an additional task for determining a configuration solution; and input the additional prompts into an additional large language model to obtain an additional response including a specified configuration solution for resolving the identified image quality problem.

[0145] Example 34 provides the device described in Example 33, wherein the prompt also includes the role of a multimodal large language model, specifying that the multimodal large language model is an expert image quality engineer.

[0146] Example 35 provides the device described in Example 33 or 34, wherein the task includes one or more types of possible image quality problems.

[0147] Example 36 provides a device described in any of Examples 33-35, wherein the task includes one or more features of possible image quality problems.

[0148] Example 37 provides the device described in Example 35 or 36, wherein possible image quality issues include lighting conditions, exposure, color balance, tone mapping, sharpness, and image noise.

[0149] Example 38 provides a device described in any of Examples 33-37, wherein the task includes instructions for comparing an image and a reference target image.

[0150] Example 39 provides a device described in any of Examples 33-38, wherein the task includes one or more of aesthetic preferences and image quality standards.

[0151] Example 40 provides any of the devices described in Examples 33-39, wherein formatting a query based on an identified image quality problem includes restating the identified image quality problem as a question.

[0152] Example 41 provides a device described in any of Examples 33-40, wherein the formatted query based on the identified image quality problem includes attaching one or more of aesthetic preferences and image quality criteria to the identified image quality problem.

[0153] Example 42 provides a device from any of Examples 33-41, wherein retrieving context from a vector database includes retrieving multiple embeddings of camera configuration knowledge that most closely matches the query embedding.

[0154] Example 43 provides a device described in any of Examples 33-42, wherein the camera configuration knowledge includes one or more of the following: camera parameter names, camera parameter value ranges for camera parameter names, camera functionality, image processing algorithm specifications, image quality tool methods, camera manuals, register configuration, and firmware configuration code.

[0155] Example 44 provides any of the devices described in Examples 33-43, wherein additional hints include the role of a multimodal large language model, specifying that the multimodal large language model is an expert image processing engineer.

[0156] Example 45 provides a device described in any of Examples 33-44, wherein the additional task includes instructions for determining a configuration solution for a box in an image processing pipeline.

[0157] Example 46 provides any of the devices described in Examples 33-45, wherein an additional task includes outputting instructions for configuring the solution in the form of one or more operations.

[0158] Example 47 provides a device described in any of Examples 33-46, wherein an additional task includes outputting instructions to write one or more register values ​​in the form of one or more register values ​​and one or more register addresses.

[0159] Example 48 provides any of the devices described in Examples 33-47, wherein an additional task includes outputting instructions for configuring the solution in the form of one or more application programming interface function calls.

[0160] Example A provides an apparatus that includes components for performing or being used to perform any of the methods provided in Examples 1-16 and the methods / processes described herein.

[0161] Example B provides the analysis phase and solution generation phase as described in this article.

[0162] Example C provides the analysis phase as described in this article.

[0163] Example D provides the solution generation phase as described in this article.

[0164] Variations and other annotations Although reference Figure 2-7 The operations of the example methods shown and described are illustrated as each one and occur once in a specific order; however, it will be appreciated that the operations can be performed in any suitable order and repeated as needed. Furthermore, one or more operations can be performed in parallel. Figure 2-7 The operations shown can be combined or may include more or less detail than described.

[0165] The foregoing description of the illustrated embodiments of this disclosure (including those described in the abstract) is not intended to be exhaustive or to limit this disclosure to the precise forms disclosed. While specific implementations and examples of this disclosure have been described herein for illustrative purposes, various equivalent modifications are possible within the scope of this disclosure, as will be recognized by those skilled in the art. These modifications can be made to this disclosure based on the detailed description above.

[0166] For illustrative purposes, specific figures, materials, and configurations have been set forth to provide a thorough understanding of the illustrative implementation. However, it will be apparent to those skilled in the art that this disclosure may be practiced without specific details, and / or may be practiced solely through some of the described aspects. In other instances, well-known features have been omitted or simplified to avoid obscuring the illustrative implementation.

[0167] Furthermore, reference has been made to the accompanying drawings, which form part of this specification, in which illustrative embodiments that may be implemented are shown. It should be understood that other embodiments may be utilized, and structural or logical changes may be made without departing from the scope of this disclosure. Therefore, the following detailed description should not be construed as limiting.

[0168] Various operations can be described as multiple discrete actions or operations in a manner most conducive to understanding the disclosed subject matter. However, the order of description should not be construed as implying that these operations are necessarily sequentially related. In particular, these operations may not be performed in the order presented. The described operations may be performed in a different order than the described embodiments. In other embodiments, various additional operations may be performed, or the described operations may be omitted.

[0169] For the purposes of this disclosure, the phrase "A or B" or the phrase "A and / or B" means (A), (B), or (A and B). For the purposes of this disclosure, the phrase "A, B, or C" or the phrase "A, B, and / or C" means (A), (B), (C), (A and B), (A and C), (B and C), or (A, B, and C). When referring to a measurement range, the term "between" includes both ends of the measurement range.

[0170] The description uses the phrases "in one embodiment" or "in an embodiment," each of which can refer to one or more of the same or different embodiments. The terms "comprising," "including," "having," etc., used with respect to embodiments of this disclosure are synonymous. This disclosure may use perspective-based descriptions, such as "above," "below," "top," "bottom," and "side," to interpret various features of the drawings; however, these terms are merely for ease of discussion and do not imply a desired or required direction. The drawings are not necessarily drawn to scale. Unless otherwise stated, the use of adjectives such as "first," "second," and "third," etc., describing a common order of objects merely indicates that different instances of the same object are referenced and does not imply that the objects so described must be in a given order in time, space, hierarchy, or any other manner.

[0171] In the following detailed description, various aspects of the illustrative implementation will be described using terms that are typically used by those skilled in the art to communicate the substance of their work to others skilled in the art.

[0172] The terms “substantially,” “close to,” “approximately,” “near,” and “about” generally refer to a value within + / - 20% of the target value, as described herein or known in the art. Similarly, terms indicating the orientation of various elements (e.g., “coplanar,” “perpendicular,” “orthogonal,” “parallel,” or any other angle between elements) generally refer to a value within + / - 5–20% of the target value, as described herein or known in the art.

[0173] Furthermore, the terms “comprise”, “comprising”, “include”, “including”, “have”, or any other variations thereof are intended to cover non-exclusive inclusion. For example, a method, process, or apparatus that includes a list of elements is not necessarily limited to those elements, but may include other elements not expressly listed or inherent to the method, process, or apparatus. Additionally, the term “or” refers to an inclusive “or”, not an exclusive “or”.

[0174] Each of the systems, methods, and apparatuses disclosed herein has several innovative aspects, none of which alone is responsible for all the desired properties disclosed herein. Details of one or more implementations of the subject matter described herein are set forth in the specification and accompanying drawings.

Claims

1. A method comprising: Generate prompts for a task that includes an image and identifies image quality issues present in that image; The prompts are input into a large multimodal language model to obtain responses that include the identification of image quality issues; The query is formatted based on the identified image quality issues in the response; Convert the query into a query embedding; The query embedding is used to retrieve context from a vector database of embeddings containing camera configuration knowledge, the context having one or more pieces of camera configuration knowledge related to the identified image quality problem; Generate additional prompts for another task that includes the query, the context, and a determination of the configuration solution; and The additional prompts are input into another large language model to obtain an additional response, which includes a specified configuration solution to address the image quality issues of the recognition.

2. The method according to claim 1, wherein, The prompt also includes the role of the multimodal large language model, specifying that the multimodal large language model is an expert image quality engineer.

3. The method according to claim 1, wherein, The task includes one or more types of possible image quality problems.

4. The method according to claim 1, wherein, The task includes one or more characteristics of potential image quality problems.

5. The method according to claim 3 or 4, wherein, The potential image quality issues include lighting conditions, exposure, color balance, tone mapping, sharpness, and image noise.

6. The method according to any one of claims 1-4, wherein, The task includes instructions to compare the image with a reference target image.

7. The method according to any one of claims 1-4, wherein, The task includes one or more of aesthetic preferences and image quality standards.

8. The method according to any one of claims 1-4, wherein, Formatting the query based on the identified image quality issues includes: The image quality problem identified is rephrased as a problem.

9. The method according to any one of claims 1-4, wherein, Formatting the query based on the identified image quality issues includes: One or more of aesthetic preferences and image quality standards are attached to the identified image quality problem.

10. The method according to any one of claims 1-4, wherein, Retrieving the context from the vector database includes: Retrieve multiple embeddings of the camera configuration knowledge that most closely match the query embedding.

11. The method according to any one of claims 1-4, wherein, The camera configuration knowledge includes one or more of the following: camera parameter name, camera parameter value range of the camera parameter name, camera functionality, image processing algorithm specifications, image quality tools and methods, camera manual, register configuration, and firmware configuration code.

12. The method according to any one of claims 1-4, wherein, The additional hints also include the role of the multimodal large language model, which specifies that the multimodal large language model is an expert image processing engineer.

13. The method according to any one of claims 1-4, wherein, The additional task includes instructions for determining the configuration solution for blocks in the image processing pipeline.

14. The method according to any one of claims 1-4, wherein, The additional task includes outputting instructions for taking one or more actions to implement the configuration solution.

15. The method according to any one of claims 1-4, wherein, The additional task includes instructions for outputting the configuration solution in the form of one or more register values ​​and one or more register addresses to write the one or more register values.

16. The method according to any one of claims 1-4, wherein, The additional task includes outputting instructions for the configuration solution in the form of one or more application programming interface function calls.

17. A non-transitory computer-readable medium storing one or more instructions, which, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1-16.

18. An apparatus comprising: One or more processors; and A non-transitory computer-readable medium storing one or more instructions, which, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1-16.