A test-time learning method for visual language large model space reasoning
Patent Information
- Application Number
- CN202610177918.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-07
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-02-07
AI Technical Summary
[0010]本发明公开了一种用于视觉语言大模型空间推理的测试时学习方法,旨在解决现有视觉语言模型在定量空间推理任务中鲁棒性差、预测数值不符合几何逻辑一致性、以及难以在线自适应特定场景的问题
[0053] 1) Strong logical self-consistency: This invention, through spatial query enhancement and Pythagorean theorem constraints, forces the visual language model to follow the geometric laws of the physical world during reasoning, effectively eliminating the random "illusion" phenomenon in numerical generation.
Smart Images

Figure CN122021775B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary field of computer vision and natural language processing, specifically to a technique for enhancing the spatial reasoning ability of large visual language models, and particularly to a test-time learning method for spatial reasoning in large visual language models. Background Technology
[0002] With the rapid development of deep learning technology, Vision-Language Models (VLMs), through pre-training on large-scale image-text pairs, have demonstrated powerful cross-modal understanding and generation capabilities, and are widely used in fields such as visual question answering, image description, and natural language instruction following. However, despite the breakthroughs achieved by existing visual language models in qualitative description and semantic recognition, they still face significant challenges in handling quantitative spatial reasoning tasks that require precise quantification calculations (such as measuring slant distance, vertical distance, and horizontal distance between objects).
[0003] First, existing models often exhibit logical inconsistencies when generating numerical results. Specifically, when predicting multiple dimensions with geometrical coupling in the same visual scene, the results often violate the geometric constraints of the physical world. This phenomenon indicates that although the model has learned the correlation mapping between visual pixels and text values, it has not truly understood the structural logic of physical space.
[0004] Secondly, existing solutions mainly rely on full parameter fine-tuning or instruction fine-tuning. However, these methods have serious limitations:
[0005] 1) High data cost: Obtaining accurate spatial measurement annotation data requires expensive sensors or intensive manual annotation, resulting in a limited scale of high-quality quantitative inference datasets;
[0006] 2) Insufficient generalization performance: The knowledge learned during the pre-training stage is often limited to the distribution of the training set. When the model is deployed and faces long-tailed distribution scenes or visual perspectives that have never been seen before, it is prone to serious "illusion" phenomena, and the output values are completely different from the visual reality.
[0007] 3) Static inference limitations: Traditional models have frozen parameters during the test inference phase, making it impossible to learn and evolve online in real time based on the specific scenario of the current input (self-adaptation), resulting in a lack of flexibility when dealing with dynamic environments.
[0008] Therefore, how to use prior geometric knowledge of the physical world as unsupervised constraints without additional manual annotation to guide the model to discover inference conflicts in real time during the testing phase and drive the model parameters to dynamically self-correct, thereby improving the accuracy and logical consistency of spatial reasoning, has become a core technical challenge that urgently needs to be solved in the field of visual language understanding. Summary of the Invention
[0009] (1) Technical problems to be solved
[0010] This invention discloses a test-time learning method for spatial reasoning of large visual language models, aiming to solve the problems of poor robustness of existing visual language models in quantitative spatial reasoning tasks, inconsistent predicted values with geometric logic, and difficulty in adapting to specific scenarios online.
[0011] (2) Technical solution
[0012] This invention discloses a test-time learning method for large-scale spatial reasoning in visual language models, comprising the following steps:
[0013] Step 1: Obtain the input image and the original query, and expand the original query into a set of auxiliary queries that satisfy geometric coupling relationships through query enhancement strategies;
[0014] Step 2: Input the auxiliary distance query into the pre-trained model to obtain the initial prediction result, use geometric constraints to verify the consistency of the initial prediction result, and convert the verified values into a structured pseudo-label distribution.
[0015] Step 3: Using the pseudo-labels as the optimization objective, construct an optimization function that includes geometric consistency loss;
[0016] Step 4: Dynamically update specific parameters of the visual language model during the inference phase by using an optimization function that minimizes the geometric consistency loss.
[0017] Furthermore, the specific steps of step 1 are as follows:
[0018] Step 101: Identify the target attributes in the original query. The target attribute Including slant distance Vertical distance or horizontal distance One of them;
[0019] Step 102: Generate the target attribute based on template transformation. An enhanced query set is constructed using auxiliary queries with geometric coupling relationships, where the geometric coupling relationship is derived from the Pythagorean theorem formula. definition.
[0020] Furthermore, step 101 specifically involves the following steps:
[0021] Step 10101: Perform entity recognition on the original query text to locate the starting point object entity to be measured in the input image. With the endpoint object entity ;
[0022] Step 10102: Extract spatial measurement keywords from the original query. The keywords are selected from one of "slope distance", "vertical distance", and "horizontal distance".
[0023] Step 10103: Perform semantic mapping based on the extracted keywords to categorize the original query into target attributes.
[0024] .
[0025] Furthermore, step 102 specifically includes the following steps:
[0026] Step 10201, based on the identified target attributes Retrieve two complementary dimension templates from a pre-defined geometric relation template library; if the original attribute is slope distance... Then retrieve the vertical distance. Template and horizontal distance template;
[0027] Step 10202, transfer the entity starting point object entity. With the endpoint object entity The templates for the two complementary dimensions retrieved are populated to generate corresponding auxiliary query text, which, together with the original query, constitute an enhanced query set. .
[0028] Furthermore, step 2 specifically includes the following steps:
[0029] Step 201: Compare the input image with the enhanced query set. Each query in the process is paired up to form multiple input pairs, which are then input into the pre-trained visual language model (VLM) to obtain the corresponding multiple sets of raw numerical prediction results.
[0030] Step 202: Establish an adaptive geometric triggering mechanism, using any two dimension values from the prediction results to calculate the reference value for the third dimension. and calculate the reference value. With the corresponding original predicted value The relative error between them;
[0031] Step 203: Determine whether the relative error is less than a preset tolerance threshold. If it is less, then determine the geometric reference value. The reliable signal is then converted into a structured pseudo-label distribution using word segmentation and serialization operations.
[0032] Furthermore, step 202 specifically includes the following steps:
[0033] Step 20201: Calculate the target attribute using the Pythagorean theorem. Geometric reference value When the target attribute Slope distance hour, ;in and The vertical and horizontal distances predicted by the model;
[0034] Step 20202: Calculate the original predicted values of the model. Compared with reference value relative error between .
[0035] Furthermore, step 203 specifically includes the following steps:
[0036] Step 20301: Determine the reference value as a reliable signal. It is parsed into a string sequence consisting of integer digits, decimal places, fractional digits, and units;
[0037] Step 20302: Utilize the preset word segmenter Perform mapping operations: This maps the continuous numerical space into a discrete sequence of word segmentation indices, where... Segmenting the integer part into words, For decimal point segmentation, For decimal part segmentation, Word segmentation for measurement units, For the reference value The discrete word segmentation index sequence.
[0038] Furthermore, the specific process of step 20302 is as follows:
[0039] Step 2030201, for fixed format positions Construct single-point distributed pseudo-labels ,Right now ,in, Indicates the position index in the word segmentation sequence. Indicates the position of the decimal point. Indicates unit position, This indicates the reference word segmentation index corresponding to this position;
[0040] Step 2030202, regarding the location of the numerical content. In the word segmentation table Filtering neighborhood word segmentation sets And construct a uniform pseudo-label distribution. :
[0041]
[0042] in, Indicates integer position. Indicates the decimal position. The complete set of word segmentation tables pre-set for the model, This is a subset of candidate word segments centered on the reference word segmentation. This indicates the total number of word segments in the subset. This is the word segmentation index in the vocabulary.
[0043] Furthermore, step 3 specifically involves the following steps:
[0044] Extracting the model at sequence positions Predicted distribution Calculate the geometric consistency loss covering the entire sequence. :
[0045]
[0046] in, Indicates the model's predicted location For word segmentation The probability distribution; For indicator functions, when index The value is 1 if it belongs to a numeric content position, and 0 otherwise. The negative likelihood penalty coefficient is preset. This represents the set of remaining word segments in the vocabulary after excluding neighboring word segments, and is used to generate numerical values to penalize physical logic errors. Indicates an integer position; Indicates the position of the decimal point; Indicates the decimal position; Indicates unit location; The complete set of word segmentation tables pre-set for the model; This represents the distribution of pseudo-labels.
[0047] Furthermore, step 4 specifically involves the following steps:
[0048] Step 401: Freeze the backbone network parameters of the visual language model, and only adjust the weights of the low-rank adaptive (LoRA) module. Set to trainable state;
[0049] Step 402, for each test sample, apply geometric consistency loss. Perform at least one gradient descent iteration, and the update formula is:
[0050]
[0051] in, The preset learning rate, This indicates the loss function with respect to the weight parameters. The gradient is used to update parameters online so that the model output is geometrically consistent.
[0052] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0053] 1) Strong logical self-consistency: This invention, through spatial query enhancement and Pythagorean theorem constraints, forces the visual language model to follow the geometric laws of the physical world during reasoning, effectively eliminating the random "illusion" phenomenon in numerical generation.
[0054] 2) No labeled data required: Adopting an adaptive geometric triggering mechanism, using geometric consistency as a natural supervision signal, it realizes test-time learning in an unsupervised environment and reduces data acquisition costs.
[0055] 3) Online scene adaptation: By dynamically updating the LoRA parameters, the model can be optimized in real time for each specific test sample, which significantly improves the spatial reasoning accuracy and generalization ability of the model in unseen scenes.
[0056] 4) Structured error correction mechanism: The introduction of negative likelihood loss term and structured pseudo-label distribution not only optimizes numerical accuracy, but also ensures the standardization of output format (such as decimal point and unit) through hard constraints. Attached Figure Description
[0057] Figure 1 This is the overall flowchart of the present invention.
[0058] Figure 2 This is a flowchart of the numerical serialization into a structured pseudo-label distribution in step 2 of the present invention. Detailed Implementation
[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0060] For the sake of clarity and reference, the technical terms, abbreviations, or acronyms used below are summarized and explained as follows:
[0061] VLM (Vision-Language Model): A large-scale visual language model, a deep learning model capable of processing both image and text inputs simultaneously.
[0062] TTL (Test-Time Learning): This refers to the process where a model adjusts its parameters in real time based on the current input samples during the inference phase.
[0063] LoRA (Low-Rank Adaptation): A low-rank adaptive method, an efficient way to fine-tune a model with a large number of parameters, which replaces the full parameter update by training a low-rank matrix.
[0064] Tokenizer: A tool that converts text strings into numerical indices (Token IDs) that can be recognized by the model.
[0065] refer to Figure 1 This invention proposes a test-time learning method for large-scale spatial reasoning in visual language, specifically including the following steps:
[0066] Step 1: Obtain the input image and the original query, and expand the original query into a set of auxiliary queries that satisfy geometric coupling relationships through query enhancement strategies;
[0067] Step 2: Input the auxiliary distance query into the pre-trained model to obtain the initial prediction result, use geometric constraints to verify the consistency of the initial prediction result, and convert the verified values into a structured pseudo-label distribution.
[0068] Step 3: Using the pseudo-labels as the optimization objective, construct an optimization function that includes geometric consistency loss;
[0069] Step 4: Dynamically update specific parameters of the visual language model during the inference phase by using an optimization function that minimizes the geometric consistency loss.
[0070] The following example illustrates a test-time learning method for large-scale spatial reasoning in visual language:
[0071] Step 1: Obtain the input image and the original query. Expand the original query into a set of auxiliary queries that satisfy geometric coupling relationships using a query enhancement strategy. Specifically, this includes the following algorithm flow:
[0072] Step 101: Identify the target attributes in the original query. The target attribute Selected from slant distance Vertical distance or horizontal distance One of them;
[0073] Specifically, step 101 includes the following sub-steps:
[0074] Step 10101: Perform entity recognition on the original query text to locate the starting point object entity to be measured in the input image. With the endpoint object entity For example, if the original query is "What is the slant distance between the drone in the picture and the charging station?", the starting point is identified as "drone" and the ending point as "charging station".
[0075] Step 10102: Extract spatial measurement keywords from the original query. The keywords are selected from one of "slope distance", "vertical distance", and "horizontal distance".
[0076] Step 10103: Perform semantic mapping based on the extracted keywords to categorize the original query into target attributes.
[0077] .
[0078] Step 102: Generate the target attribute based on template transformation. An enhanced query set is constructed using auxiliary queries with geometric coupling relationships, where the geometric coupling relationship is derived from the Pythagorean theorem formula. definition;
[0079] Specifically, step 102 includes the following sub-steps:
[0080] Step 10201, based on the identified target attributes Retrieve two complementary dimension templates from a pre-defined geometric relation template library; if the original attribute is slope distance... Then retrieve the vertical distance. Template and horizontal distance template;
[0081] Step 10202, transfer the entity starting point object entity. With the endpoint object entity The templates for the two complementary dimensions retrieved are populated to generate corresponding auxiliary query text, which, together with the original query, constitute an enhanced query set. .
[0082] Step 2: Input the auxiliary distance query into the pre-trained model to obtain the initial prediction results. Use geometric constraints to verify the consistency of the initial prediction results, and convert the verified values into a structured pseudo-label distribution, referencing... Figure 2 ;
[0083] Specifically, step 2 includes the following algorithm flow:
[0084] Step 201: Compare the input image with the enhanced query set. Each query in the process is paired up to form multiple input pairs, which are then input into a pre-trained visual language model (VLM) to obtain multiple sets of original numerical prediction results.
[0085] Step 202: Establish an adaptive geometric triggering mechanism, using any two dimension values from the prediction results to calculate the reference value for the third dimension. and calculate the reference value. With the corresponding original predicted value The relative error between them;
[0086] Specifically, step 202 includes the following sub-steps:
[0087] Step 20201: Calculate the target attribute using the Pythagorean theorem. Geometric reference value When the target attribute Slope distance hour, ;in and The vertical and horizontal distances predicted by the model;
[0088] Step 20202: Calculate the original predicted values of the model. Compared with reference value relative error between .
[0089] Step 203: Determine whether the relative error is less than a preset tolerance threshold. If it is less than, then the geometric reference value is determined. The reliable signal is then converted into a structured pseudo-label distribution using word segmentation and serialization operations.
[0090] Specifically, step 203 includes the following sub-steps:
[0091] Step 20301: Determine the reference value as a reliable signal. The value 5.36 meter is parsed as a string sequence consisting of integer digits, decimal places, fractional digits, and unit. For example, the value 5.36 meter is parsed as the string sequence ["5", ".", "36", "meter"].
[0092] Step 20302: Utilize the preset word segmenter Perform mapping operations: This maps the continuous numerical space into a discrete sequence of word segmentation indices, where... Segmenting the integer part into words, For decimal point segmentation, For decimal part segmentation, Word segmentation for measurement units, For the reference value The discrete word segmentation index sequence;
[0093] Specifically, the process of constructing the structured pseudo-label distribution in step 20302 is as follows:
[0094] Step 2030201, for fixed format positions Construct single-point distributed pseudo-labels ,Right now ,in, Indicates the position index in the word segmentation sequence. Indicates the position of the decimal point. Indicates unit position, This indicates the reference word segmentation index corresponding to this position;
[0095] Step 2030202, regarding the location of the numerical content. In the word segmentation table Filtering neighborhood word segmentation sets (e.g., a subset of tokens adjacent to the target value 5), and construct a uniform pseudo-label distribution. :
[0096]
[0097] in, Indicates integer position. Indicates the decimal position. The complete set of word segmentation tables pre-set for the model, This is a subset of candidate word segments centered on the reference word segmentation. This indicates the total number of word segments in the subset. This is the word segmentation index in the vocabulary.
[0098] Step 3: Using the pseudo-labels as the optimization objective, construct an optimization function that includes geometric consistency loss;
[0099] Specifically, the model extracts information from sequence positions. Predicted distribution Calculate the geometric consistency loss covering the entire sequence. :
[0100]
[0101] in, Indicates the model's predicted location For word segmentation The probability distribution; For indicator functions, when index The value is 1 if it belongs to a numeric content position, and 0 otherwise. The negative likelihood penalty coefficient is preset. This represents the set of remaining word segments in the vocabulary after excluding neighboring word segments, and is used to generate numerical values to penalize physical logic errors. Indicates an integer position; Indicates the position of the decimal point; Indicates the decimal position; Indicates unit location; The complete set of word segmentation tables pre-set for the model; This represents the distribution of pseudo-labels. For example, the formula guides the model to generate the correct number (5) through the left term, and severely punishes the generation of logically incorrect numbers (such as 9) through the right term (negative likelihood term), ensuring that the generated values are geometrically and logically closed-loop.
[0102] Step 4: Dynamically update specific parameters of the visual language model during the inference phase by using an optimization function that minimizes the geometric consistency loss;
[0103] Specifically, the algorithm flow includes the following:
[0104] Step 401: Freeze the backbone network parameters of the visual language model, and only adjust the weights of the low-rank adaptive (LoRA) module. Set to trainable state;
[0105] Step 402, for each test sample, apply geometric consistency loss. Perform at least one gradient descent iteration, and the update formula is:
[0106]
[0107] in, The preset learning rate, This indicates the loss function with respect to the weight parameters.
[0108] The gradient is used to update the parameters online so that the model output is consistent in the geometric dimension (for example, for a specific viewpoint, the model is revised from the initial prediction of 5.36 meters to 5.02 meters, making it more consistent with the true geometric result under the Pythagorean theorem constraint).
[0109] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0110] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0111] The technical features of the above embodiments can be combined arbitrarily. Furthermore, the numbering of each step is not intended to constrain the order of the steps; their order is permissible as long as there are no strict constraints on the sequence. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, any combination of these technical features that does not contradict each other should be considered within the scope of this specification.
[0112] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A test-time learning method for visual language large model space reasoning, characterized in that: Includes the following steps: Step 1: Obtain the input image and the original query, and expand the original query into a set of auxiliary queries that satisfy geometric coupling relationships through query enhancement strategies; Step 2: Input the auxiliary query into the pre-trained model to obtain the initial prediction result, use geometric constraints to verify the consistency of the initial prediction result, and convert the verified values into a structured pseudo-label distribution. Step 2 specifically includes the following steps: Step 201: Pair the input image with each query in the enhanced query set to form multiple input pairs and input them into the pre-trained visual language model (VLM) to obtain the corresponding multiple sets of original numerical prediction results. Step 202: Establish an adaptive geometric triggering mechanism, use any two dimension values in the prediction result to deduce the reference value of the third dimension, and calculate the relative error between the reference value and the corresponding original prediction value; Step 203: Determine whether the relative error is less than a preset tolerance threshold. If it is less, determine that the geometric reference value is a reliable signal, and use word segmentation and serialization operations to convert the reliable signal into a structured pseudo-label distribution. Step 3: Using the pseudo-labels as the optimization objective, construct an optimization function that includes geometric consistency loss; The specific steps of step 3 are as follows: extracting the model at the sequence position of the prediction distribution , computing the geometric consistency loss covering the whole sequence, computing the geometric consistency loss covering the whole sequence : ; in, Indicates the model's predicted location For word segmentation The probability distribution; For indicator functions, when index The value is 1 if it belongs to a numeric content position, and 0 otherwise. The negative likelihood penalty coefficient is preset. This represents the set of remaining word segments in the vocabulary after excluding neighboring word segments, and is used to generate numerical values to penalize physical logic errors. Indicates an integer position; Indicates the position of the decimal point; Indicates the decimal position; Indicates the unit location; The complete set of word segmentation tables pre-set for the model; Indicates the distribution of pseudo-labels; Step 4: Dynamically update specific parameters of the visual language model during the inference phase by using an optimization function that minimizes the geometric consistency loss.
2. The test-time learning method for large-scale spatial reasoning in visual language according to claim 1, characterized in that: The specific steps of step 1 are as follows: Step 101: Identify the target attributes in the original query. The target attribute Including slant distance Vertical distance or horizontal distance One of them; Step 102: Generate the target attribute based on template transformation. An enhanced query set is constructed using auxiliary queries with geometric coupling relationships, where the geometric coupling relationship is derived from the Pythagorean theorem. definition.
3. The test-time learning method for large-scale spatial reasoning in visual language according to claim 2, characterized in that: The specific steps of step 101 are as follows: Step 10101: Perform entity recognition on the original query text to locate the starting point object entity to be measured in the input image. With the endpoint object entity ; Step 10102: Extract spatial measurement keywords from the original query. The keywords are selected from one of "slope distance", "vertical distance", and "horizontal distance". Step 10103: Perform semantic mapping based on the extracted keywords to categorize the original query into target attributes. .
4. The test-time learning method for large-scale spatial reasoning in visual language according to claim 3, characterized in that: The specific steps of step 102 are as follows: Step 10201, based on the identified target attributes Retrieve two complementary dimension templates from a pre-defined geometric relation template library; if the original attribute is slope distance... Then retrieve the vertical distance. Template and horizontal distance template; Step 10202, solidify the starting object With the endpoint object entity The templates for the two complementary dimensions retrieved are populated to generate corresponding auxiliary query text, which, together with the original query, constitute an enhanced query set. .
5. The test-time learning method for large-scale spatial reasoning in visual language according to claim 4, characterized in that: Step 202 specifically includes the following steps: Step 20201: Calculate the target attribute using the Pythagorean theorem. Geometric reference value When the target attribute Slope distance hour, ;in and The vertical and horizontal distances predicted by the model; Step 20202: Calculate the original predicted values of the model. Compared with reference value relative error between .
6. The test-time learning method for large-scale spatial reasoning in visual language according to claim 4, characterized in that: Step 203 specifically includes the following steps: Step 20301: Determine the reference value as a reliable signal. It is parsed into a string sequence consisting of integer digits, decimal place, decimal places, and units; Step 20302: Utilize a preset word segmenter Perform mapping operations: This maps the continuous numerical space into a discrete sequence of word indexes, where... Segmenting the integer part into words, For decimal point segmentation, For decimal part segmentation, Word segmentation for measurement units, For the reference value The discrete word segmentation index sequence.
7. The test-time learning method for large-scale spatial reasoning in visual language according to claim 6, characterized in that: The specific process of step 20302 is as follows: Step 2030201, for fixed format positions Construct single-point distributed pseudo-labels ,Right now ,in, Indicates the position index in the word segmentation sequence. Indicates the position of the decimal point. Indicates unit position, This indicates the reference word segmentation index corresponding to this position; Step 2030202, regarding the location of the numerical content. In the word segmentation table Filtering neighborhood word segmentation sets And construct a uniform pseudo-label distribution. : ; in, Indicates integer position. Indicates the decimal position. The complete set of word segmentation tables pre-set for the model, This is a subset of candidate word segments centered on the reference word segmentation. This indicates the total number of word segments in the subset. This is the word segmentation index in the vocabulary.
8. The test-time learning method for large-scale spatial reasoning in visual language according to claim 7, characterized in that: The specific steps of step 4 are as follows: Step 401: Freeze the backbone network parameters of the visual language model, and only adjust the weights of the low-rank adaptive LoRA module. Set to trainable state; Step 402, for each test sample, apply geometric consistency loss. Perform at least one gradient descent iteration, and the update formula is: ; in, The preset learning rate, This indicates the loss function with respect to the weight parameters. The gradient is used to update parameters online so that the model output is geometrically consistent.
Citation Information
Patent Citations
Intelligent driving target detection method based on vision-language large model fine tuning
CN118155182A
Systems and methods for hybrid artificial intelligence enhancement and optimization
WO2025064722A1