Geometric constraint enhanced large language model space cognitive optimization method and system

By building a structured training set and polar coordinate encoder, combining path memory and planning modules, optimizing the spatial cognitive ability of large language models, the problems of spatial relationship reasoning errors and insufficient dynamic spatial modeling capabilities in the existing technology are solved, and higher accuracy and reliability are achieved.

CN120471180AInactive Publication Date: 2025-08-12SHEYUE FUTURE INFORMATION TECHNOLOGY (HANGZHOU) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510971923.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-08-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing large language models have spatial relationship inference errors, weak dynamic spatial evolution modeling capabilities in spatial cognitive tasks, and lack effective mapping of Euclidean spatial metric relationships and unstructured texts, resulting in limited applications in fields such as intelligent navigation, virtual reality interaction and building information modeling.

Method used

Build a structured multimodal training set that integrates orientation instructions, maze path planning and three-dimensional structure cognition, combines polar coordinate spatial encoder and path memory and planning modules, and improves the spatial cognitive ability of the model through supervision fine-tuning and direct preference optimization.

Benefits of technology

It significantly improves the accuracy of the model in azimuth analysis, path planning and three-dimensional reconstruction tasks, solves the error accumulation problem of traditional models in spatial cognitive tasks, and achieves higher geometric consistency and engineering reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471180A_ABST
    Figure CN120471180A_ABST
Patent Text Reader

Abstract

The invention discloses a geometric constraint enhanced large language model space cognition optimization method and system. A geometric constraint enhanced large language model is adopted to determine the orientation and space relation in a task; wherein the geometric constraint enhancement-based large language model is specifically implemented by the following steps of: constructing a structured multi-modal training set fusing an orientation instruction, labyrinth path planning and three-dimensional structure cognition; a polar coordinate space encoder and a path memorizing and planning module are combined in a large language model, the polar coordinate space encoder is used for dynamically mapping azimuth description in a natural language into computable geometric vectors, and accurate association between text semantics and space coordinates is established; the path memorizing and planning module is used for realizing time sequence memorizing and dynamic re-planning of a path instruction; performing supervision fine tuning and direct preference optimization based on a staged parameter updating strategy on the pre-trained large language model; and quantitative evaluation is carried out by adopting a multi-dimensional test framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a method and system for optimizing the spatial cognition of a large language model with enhanced geometric constraints, in particular, a training method and system architecture for improving the spatial reasoning ability of a large language model by integrating supervised fine-tuning (SFT), reinforcement learning with human feedback (RLHF), and a spatial cognition experimental paradigm. Background Art

[0002] In recent years, large language models (LLMs) have demonstrated remarkable capabilities in text generation and comprehension in natural language processing, and their application has expanded to encompass diverse areas, including intelligent customer service, text summarization, and code generation. However, existing technical literature indicates that traditional LLM architectures exhibit systematic flaws when handling spatial cognition tasks. These flaws manifest as "spatial illusions," resulting from errors in spatial relationship reasoning, inaccurate three-dimensional scene descriptions, and weak modeling of dynamic spatial evolution. This technical bottleneck severely limits the potential of LLMs in space-intensive applications, such as intelligent navigation systems, virtual reality interaction, and Building Information Modeling (BIM).

[0003] The current mainstream LLM training paradigm relies primarily on learning statistical patterns from text corpora and lacks the ability to structure spatial information. Research has shown that when processing text containing directional prepositions, topological relationships, or dynamic motion trajectories, existing large language models have a high error rate in directional judgment. Furthermore, their accuracy lags significantly behind that of human experts in parsing multi-layered spatial nesting structures. This shortcoming stems from the lack of a spatial cognition module in existing model architectures, particularly the lack of an effective learning paradigm for the mapping mechanism between Euclidean spatial metric relationships and unstructured text.

[0004] Existing technologies attempt to improve spatial cognition through multimodal fusion methods, such as the joint vision-language training strategy employed by the CLIP architecture. However, these approaches still face significant limitations in spatial relationship reasoning: first, 2D image feature extraction struggles to capture the hierarchical structure of 3D space; second, visual-text alignment mechanisms lack modeling of projective transformation invariance. Experimental data show that on the standard SUN397 spatial relationship dataset, multimodal models achieve less improvement in orientation prediction accuracy than single-modal approaches. This confirms the technical reality that relying solely on visual feature enhancement cannot fundamentally address deficiencies in spatial cognition.

[0005] Another area of improvement focuses on knowledge graph enhancement technology, such as the entity relationship injection mechanism adopted by ERNIE 3.0. However, existing knowledge graphs tend to focus on semantic associations, and their spatial attribute annotations suffer from coarse granularity and a single dimension. For example, 78.6% of entries in the Wikidata spatial relationship database contain only basic topological relationships such as "located in" and "nearby," lacking descriptions of refined spatial attributes such as direction angles, distance thresholds, and occlusion relationships. This limitation in knowledge representation means that enhanced LLMs can still produce contradictory outputs in complex spatial reasoning tasks, such as simultaneously asserting that "object A is northeast of object B" and "object B blocks the view to the east of object A."

[0006] At the model architecture level, while the Transformer's self-attention mechanism excels at capturing sequential dependencies, it lacks a fundamental understanding of geometric properties of spatial coordinate systems, such as rotation and scale invariance. Existing position encoding schemes (such as sinusoidal encoding and relative position encoding) primarily serve to model sequential order and are ineffective at representing rigid body transformations in three-dimensional space. This architectural flaw directly leads to cumulative errors in LLMs when processing chains of spatial transformations. For example, in recursive navigation command generation tasks, the path description error of traditional models grows exponentially with the command step length.

[0007] Analysis of technical bottlenecks reveals that the root causes of the spatial cognition deficiencies of existing LLMs lie in: 1) a lack of structured spatial annotations in training data; 2) a lack of built-in geometric reasoning mechanisms in the model architecture; and 3) a lack of quantitative spatial ability metrics in the evaluation system. While specialized datasets such as SpaceQA have emerged in the field, their insufficient data diversity (covering only 2D planar layouts) and coarse annotation granularity (for binary spatial relationship judgment) make it difficult to systematically cultivate complex spatial cognition skills. Furthermore, existing research often relies on traditional NLP metrics such as accuracy, failing to establish a comprehensive evaluation system that encompasses dimensions such as spatial consistency verification and geometric constraint satisfaction.

[0008] Based on this, this paper proposes a method for optimizing spatial cognition in large language models based on geometric constraint enhancement. Its core innovation lies in the construction of a specialized spatial cognition training system. Through the synergy of spatial data injection and targeted architecture optimization, it overcomes the theoretical limitations of traditional LLMs in modeling spatial relationships. Unlike existing partial improvements, this method reconstructs the spatial representation paradigm from the data source, introduces geometric prior knowledge during the model pre-training phase, and lays a foundation for spatial consistency for subsequent fine-tuning tasks. It also uses human reinforcement learning to enhance the relevant model capabilities. Summary of the Invention

[0009] This paper addresses the core technical flaw of existing language models, which suffer from insufficient spatial cognition. It proposes a cognitive enhancement framework that integrates structured spatial representation with reinforcement learning. By constructing multiple specialized training sets, employing a polar coordinate embedding algorithm, and reinforcing chain thinking, this framework dynamically transforms orientation relationships described in natural language into computable geometric constraints. Furthermore, it establishes a cognitive connection between textual instructions and geometric space, significantly improving the model's accuracy in tasks such as orientation parsing and path planning.

[0010] The technical solution adopted in the present invention is as follows:

[0011] A method for optimizing spatial cognition using a large language model augmented with geometric constraints, employing a large language model augmented with geometric constraints to determine the orientation and spatial relationships in tasks including, but not limited to, positioning, navigation, orientation resolution, and trajectory planning. The specific implementation of the large language model augmented with geometric constraints includes the following:

[0012] Construct a structured multimodal training set that integrates orientation instructions, maze path planning, and 3D structure recognition;

[0013] The large language model is combined with a polar coordinate space encoder and a path memory and planning module, wherein the polar coordinate space encoder is used to dynamically map the orientation description in natural language into a computable geometric vector, establishing a precise association between text semantics and spatial coordinates; the path memory and planning module is used to implement temporal memory and dynamic replanning of path instructions;

[0014] Supervised fine-tuning of pre-trained large language models using a phased parameter update strategy and direct preference optimization;

[0015] A multi-dimensional testing framework is used for quantitative evaluation.

[0016] In the above technical solution, further, the structured multimodal training set includes the following:

[0017] (1) Extract 50,000 polar coordinate position description instructions from geographic navigation corpus, with the format of "the target is located at direction meters", covering and Meter range, and normalize the distance value;

[0018] (2) Generate 2000 topological mazes based on random Prim algorithm, with the number of nodes After obtaining the optimal path through breadth-first search, the path deviation noise and inefficiency noise are injected to construct 30,000 sets of comparison data; in one embodiment of the present invention, the path deviation noise is to randomly replace 10% of the steering actions with invalid operations, and the inefficiency noise is to extend the path to the optimal solution. times;

[0019] (3) Convert the MRT question bank of the mental rotation test into text-matrix mapping data, where the input is "rotation around the axis Degree" instruction, the output is the corresponding Rotation matrix, with an error tolerance of Frobenius norm .

[0020] Furthermore, the polar coordinate space encoder encodes the rigid constraints of the geometric transformation into the model parameter space based on the polar coordinate bilinear projection and dynamic normalization algorithm, which specifically includes the following:

[0021] (S2.1) Polar coordinate parameter extraction: A parameter parsing module based on regular rules extracts the azimuth angle from the input text and distance ,in Convert to continuous values using the degrees to radians formula ,distance Normalized according to the maximum detection range of 50 meters ;

[0022] (S2.2) Bilinear Projection Modeling: Construct a bilinear transformation layer to implement the mapping of text embedding to geometric space. Specifically, the input text is encoded into a 4096-dimensional vector by the BERT model and then transformed into a vector by a trainable matrix. Dimensionality reduction, GELU activation function and Dimensionality increase, the output contains The geometric vector of ,preserves semantic continuity through manifold learning;

[0023] (S2.3) Dynamic normalization mechanism: A distance-sensitive factor is introduced to optimize gradient propagation, that is, a smoothing term based on the standard deviation of the distance within the batch is applied to the normalized distance.

[0024] Furthermore, the path memory and planning module includes:

[0025] (S3.1) Design a gated recurrent unit (GRU) and construct a three-layer GRU network to record movement trajectories. Its input is a polar coordinate vector. The output of each GRU layer serves as the input of the next layer, which outputs the planned trajectory.

[0026] (S3.2) Real-time path replanning: When the deviation between the detected position and the planned position exceeds a preset threshold, the A* algorithm is triggered to search for an alternative path. The cost function is composed of the actual movement cost, the steering penalty term, and the heuristic function. After the new path is injected into the GRU, the trajectory is updated through the state reset mechanism.

[0027] Furthermore, the supervised fine-tuning (SFT) adopts a phased parameter update strategy to freeze the high-level layers of the base model while enhancing spatial cognition capabilities in a targeted manner, including:

[0028] (S4.1) Hierarchical training: Freeze the last two layers of the large language model to preserve general semantic knowledge, and unfreeze the polar coordinate projection matrix Perform full training with the GRU network and insert the adapter module;

[0029] (S4.2) Calculate the multi-objective loss function for supervised fine-tuning training, where the total loss is composed of orientation loss, path cross entropy loss, and structure contrast loss with different weights.

[0030] Furthermore, a Direct Preference Optimization (DPO) mechanism is used to optimize the path planning strategy under limited parameter updates to ensure that the generated path conforms to human preferences. The DPO mechanism includes:

[0031] (S5.1) Sparse parameter update strategy: Freeze the GRU and adapter parameters trained in the SFT phase and update them through binary masks , realize only polar coordinate projection matrix Implementation Sparse update, the update formula is:

[0032]

[0033] in is the dynamically decaying learning rate , , is the preference loss function, is the Hadamard product operator;

[0034] (S5.2) The preference loss function adopts the DPO loss function:

[0035]

[0036] in is the frozen SFT model, To optimize the strategy, For input, For positive samples, is a negative sample; at the same time, a KL divergence constraint is added to prevent the strategy from deviating too far:

[0037] .

[0038] Furthermore, the establishment of a multi-dimensional testing framework includes the following:

[0039] (1) Basic orientation analysis: The geometric mapping accuracy of text instructions is quantified by the polar coordinate error rate model, covering basic capability verification in static scenarios;

[0040] (2) Dynamic path planning: In a hierarchical maze environment, the first-time planning success rate and path efficiency are calculated to evaluate the model's real-time decision-making ability in a dynamic environment;

[0041] (3) High-order structural reasoning: Based on the three-dimensional object transformation task of the Mental Rotation Test (MRT), the model's ability to abstractly understand complex spatial relationships is verified through key point matching accuracy and rotation matrix Riemann error.

[0042] A large language model spatial cognitive optimization system with enhanced geometric constraints, used to implement the method described in any of the above items.

[0043] An electronic device, comprising:

[0044] one or more processors;

[0045] a memory for storing one or more programs;

[0046] When the one or more programs are executed by the one or more processors, the one or more processors implement any of the above methods.

[0047] A computer-readable storage medium stores computer-executable instructions, wherein the instructions are used to implement any of the methods described above when executed.

[0048] The beneficial effects of the present invention are at least as follows:

[0049] This invention systematically addresses the core shortcomings of traditional large language models in spatial cognition tasks by constructing a structured spatial representation paradigm and a geometric constraint injection mechanism. Compared to existing technologies, the proposed solution achieves breakthroughs in three key areas: First, it proposes a polar coordinate bilinear projection and dynamic normalization algorithm to encode rigid constraints of geometric transformations (such as rotational covariance and translational invariance) into the model parameter space, significantly improving the mathematical rigor of positional description and 3D structure reasoning. Second, it designs a reinforcement learning strategy based on a gated path memory network and sparse parameter optimization to effectively integrate error accumulation suppression and human preferences in dynamic path planning, ensuring the physical rationality and real-time performance of complex spatial decision-making. Third, it innovatively establishes a multi-dimensional testing framework covering geometric consistency verification and dynamic complexity tolerance assessment, filling a gap in existing NLP evaluation systems for quantitative analysis of spatial attributes. These technical breakthroughs enable the model to demonstrate cognitive accuracy and engineering reliability that surpass traditional multimodal approaches in tasks such as positional parsing, path planning, and 3D reconstruction, providing a language model solution with strict geometric constraints for spatially intensive scenarios such as intelligent navigation and digital twins. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 1 is a flow chart of a large language model spatial cognition optimization method enhanced by geometric constraints in an embodiment of the present invention.

[0051] Figure 2 yes Figure 1 Flowchart of step S2 in FIG.

[0052] Figure 3 yes Figure 1 Flowchart of step S3 in FIG.

[0053] Figure 4 yes Figure 1 Flowchart of step S4 in FIG.

[0054] Figure 5 yes Figure 1 Flowchart of step S5 in FIG.

[0055] Figure 6 yes Figure 1 Flowchart of step S6 in FIG.

[0056] Figure 7 The invention is a software interface for spatial cognition optimization of a large language model with enhanced geometric constraints in an embodiment. DETAILED DESCRIPTION

[0057] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0058] like Figure 1As shown, according to a specific embodiment of the present invention, the geometric constraint-enhanced large language model spatial cognition optimization method of the present invention includes the following: (S1) Construction of a structured training set. To address the problems of scattered spatial cognition data and single annotation dimensions in the prior art, the present invention constructs a multimodal training set that integrates orientation instructions, maze path planning, and three-dimensional structure cognition. Specifically, it includes:

[0059] (1) Extract 50,000 polar coordinate position description instructions from geographic navigation corpus, with the format of "the target is located at direction meters", covering and The distance value is normalized.

[0060] (2) Generate 2000 topological mazes (number of nodes) based on random Prim algorithm ), after obtaining the optimal path through breadth-first search, inject path deviation noise (randomly replace 10% of steering actions with invalid operations) and inefficiency noise (path is extended to the optimal solution) times) to construct 30,000 sets of comparison data.

[0061] (3) Convert the Mental Rotation Test (MRT) question bank into text-matrix mapping data, where the input is "rotation around the axis Degree" instruction, the output is the corresponding Rotation matrix, with an error tolerance of Frobenius norm The subscript pred is the predicted value, and the subscript true is the true value.

[0062] (S2) Polar coordinate space encoder design, which dynamically maps the orientation description in natural language into a computable geometric vector, and establishes an accurate association between text semantics and spatial coordinates. Figure 2 , the polar coordinate space encoder includes:

[0063] (S2.1) Polar coordinate parameter extraction. Design a parameter parsing module based on regular rules to extract the azimuth angle from the text. and distance ,in Formula for converting degrees to radians

[0064]

[0065] Convert to continuous value, distance Normalized according to the maximum detection range of 50 meters

[0066]

[0067] (S2.2) Bilinear projection modeling. Construct a bilinear transformation layer to implement the mapping of text embedding to geometric space: After the input text is encoded into a 4096-dimensional vector by the model, it is transformed into a vector through a trainable matrix. Dimensionality reduction, GELU activation function and Dimensionality increase, the output contains The geometric vector

[0068]

[0069] The design preserves semantic continuity through manifold learning.

[0070] (S2.3) Dynamic normalization mechanism. Introducing distance-sensitive factors to optimize gradient propagation: applying a smoothing term to the normalized distance

[0071]

[0072] in is the standard deviation of the within-batch distance.

[0073] (S3) Path memory and planning module. This module implements the temporal memory and dynamic replanning capabilities of path instructions, solving the error accumulation problem of traditional models in long-range navigation. Figure 3 As shown, the path memory and planning module includes:

[0074] (S3.1) Gated Recurrent Unit (GRU) design. Construct a 3-layer GRU network to record movement trajectory and hidden state. , the update formula is as follows:

[0075]

[0076]

[0077]

[0078]

[0079] The input is a polar coordinate vector.

[0080] (S3.2) Real-time path replanning. When position deviation is detected:

[0081]

[0082] in Current path, For the planned path.

[0083] Trigger the A* algorithm to search for alternative paths. The cost function is:

[0084]

[0085] in Contains the actual movement cost and the turning penalty , is a vector and Angle.

[0086] is the heuristic function. The new path After injecting into GRU, the trajectory is updated through the state reset mechanism:

[0087]

[0088] (S4) Supervised Fine-tuning (SFT) strategy. Through a phased parameter update strategy, spatial cognition is enhanced while freezing the high-level layers of the base model. The supervised fine-tuning process is shown in Figure 4 and includes:

[0089] (S4.1) Hierarchical training mechanism

[0090] Freeze the last two layers of the large language model to retain general semantic knowledge and unfreeze the polar coordinate projection matrix Full training with GRU network. Insert adapter module:

[0091]

[0092] Where h is the adapter input, .

[0093] (S4.2) Calculate the multi-objective loss function for supervised fine-tuning training, where the total loss is:

[0094]

[0095] Specific definitions include:

[0096] (1) Azimuth loss:

[0097]

[0098] in is the number of samples, is the true distance of sample i, is the estimated distance output by the neural network of sample i, is the true azimuth of sample i, is the estimated azimuth angle output by the neural network for sample i.

[0099] (2) Path cross entropy:

[0100]

[0101] (3) Structural contrast loss:

[0102]

[0103] in For Samples with similar structures (such as adjacent nodes, sequence points on the same path), For Samples with dissimilar structures (such as non-adjacent nodes, sequence points on different paths), for or .

[0104] (S5) Direct Preference Optimization (DPO) mechanism. Optimizes the path planning strategy under limited parameter updates to ensure that the generated path conforms to human preferences. Figure 5 As shown, including:

[0105] (S5.1) Sparse parameter update strategy. Freeze the GRU and adapter parameters trained in the SFT phase and update them through binary masks. , realize only polar coordinate projection matrix Implementation Sparse update. The update formula is:

[0106]

[0107] in is the dynamically decaying learning rate .

[0108] (S5.2) Preference loss function, using DPO loss function:

[0109]

[0110] in is the frozen SFT model, To optimize the strategy, For input, For positive samples, is a negative sample. At the same time, a KL divergence constraint is added to prevent the strategy from deviating too far:

[0111]

[0112] (S6) Quantify the evaluation system and establish a multi-dimensional testing framework to achieve objective verification of technical effects, such as Figure 6 :

[0113] (S6.1) Azimuth resolution accuracy test. Construct 1000 sets of azimuth command-coordinate true value pairs and calculate the normalized error rate:

[0114]

[0115] (S6.2 Maze navigation efficiency evaluation. Design a 10-level dynamic maze (complexity ), measure the first planning success rate and path efficiency:

[0116]

[0117] (S6.3) Structural recognition test. Using a self-constructed MRT dataset, calculate the key point matching degree:

[0118]

[0119] (S7): 3D structural cognition verification and testing framework, which comprehensively verifies the spatial cognition ability of the technical solution and includes three levels:

[0120] (1) Basic orientation analysis: Through the polar coordinate error rate ( ) Quantify the geometric mapping accuracy of the model to text instructions, covering basic capability verification in static scenarios;

[0121] (2) Dynamic path planning: In a hierarchical maze environment (complexity C = area, number of obstacles × number of branches), the first-time planning success rate and path efficiency η are calculated to evaluate the model's real-time decision-making ability in a dynamic environment;

[0122] (3) High-order structure reasoning: A 3D object transformation task based on the Mental Rotation Test (MRT), using key point matching accuracy (Acc) and rotation matrix Riemann error ( ,

[0123] in for Rotation matrix, ), verifying the model's ability to abstractly understand complex spatial relationships.

[0124] like Figure 7 This is a schematic diagram of the spatial cognition optimization software interface designed based on the method of the present invention. Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk drives, CD-ROMs, optical storage devices, etc.) containing computer-usable program code.

[0125] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0126] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0128] The embodiments described above are merely some preferred embodiments of the present invention and are not intended to limit the present invention. Persons skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent substitution or equivalent transformation falls within the scope of protection of the present invention.

Claims

1. A method for spatial cognition optimization of large language models enhanced by geometric constraints, characterized in that: A large language model enhanced with geometric constraints is used to determine the orientation and spatial relationship in the task. The specific implementation of the large language model enhanced with geometric constraints includes the following: Construct a structured multimodal training set that integrates orientation instructions, maze path planning, and 3D structure recognition; The large language model is combined with a polar coordinate space encoder and a path memory and planning module, wherein the polar coordinate space encoder is used to dynamically map the orientation description in natural language into a computable geometric vector, establishing a precise association between text semantics and spatial coordinates; the path memory and planning module is used to implement temporal memory and dynamic replanning of path instructions; Supervised fine-tuning of pre-trained large language models using a phased parameter update strategy and direct preference optimization; A multi-dimensional testing framework is used for quantitative evaluation.

2. The geometric constraint-enhanced large language model spatial cognition optimization method according to claim 1 is characterized in that: The structured multimodal training set includes the following: (1) Extract 50,000 polar coordinate position description instructions from geographic navigation corpus, the format is "the target is located at direction meters", covering and Meter range, and normalize the distance value; (2) Generate 2000 topological mazes based on random Prim algorithm, with the number of nodes ,After obtaining the optimal path through breadth-first search, path deviation noise and ,inefficiency noise were injected to construct 30,000 sets of comparison data; (3) Convert the MRT question bank of the mental rotation test into text-matrix mapping data, where the input is "rotation around the axis Degree" instruction, the output is the corresponding Rotation matrix, with an error tolerance of Frobenius norm .

3. The method for spatial cognition optimization of large language models with enhanced geometric constraints according to claim 1, characterized in that: The polar coordinate space encoder encodes the rigid constraints of geometric transformation into the model parameter space based on polar coordinate bilinear projection and dynamic normalization algorithm, which specifically includes the following: (S2.1) Polar coordinate parameter extraction: A parameter parsing module based on regular rules extracts the azimuth angle from the input text and distance ,in Convert to continuous values using the degrees to radians formula ,distance Normalized according to the maximum detection range of 50 meters ; (S2.2) Bilinear projection modeling: Construct a bilinear transformation layer to implement the mapping of text embedding to geometric space. Specifically, the input text is encoded into a 4096-dimensional vector by the BERT model and then transformed into a vector by a trainable matrix. Dimensionality reduction, GELU activation function and Dimensionality increase, the output contains The geometric vector of ,preserves semantic continuity through manifold learning; (S2.3) Dynamic normalization mechanism: A distance-sensitive factor is introduced to optimize gradient propagation, that is, a smoothing term based on the standard deviation of the distance within the batch is applied to the normalized distance.

4. The method for spatial cognition optimization of large language models with enhanced geometric constraints according to claim 3, characterized in that: The path memory and planning module includes: (S3.1) Design a gated recurrent unit (GRU) and construct a three-layer GRU network to record movement trajectories. Its input is a polar coordinate vector. The output of each GRU layer serves as the input of the next layer, which outputs the planned trajectory. (S3.2) Real-time path replanning: When the deviation between the detected position and the planned position exceeds a preset threshold, the A* algorithm is triggered to search for an alternative path. The cost function is composed of the actual movement cost, the steering penalty term, and the heuristic function. After the new path is injected into the GRU, the trajectory is updated through the state reset mechanism.

5. The method for optimizing large language model space cognition with geometric constraint enhancement according to claim 4, characterized in that: The supervised fine-tuning (SFT) method enhances spatial cognition capabilities while freezing the high layers of the base model through a phased parameter update strategy, including: (S4.1) Hierarchical training: Freeze the last two layers of the large language model to preserve general semantic knowledge, and unfreeze the polar coordinate projection matrix Perform full training with the GRU network and insert the adapter module; (S4.2) Calculate the multi-objective loss function for supervised fine-tuning training, where the total loss is composed of orientation loss, path cross entropy loss, and structure contrast loss with different weights.

6. The method for spatial cognition optimization of large language models enhanced by geometric constraints according to claim 5, characterized in that: The Direct Preference Optimization (DPO) mechanism is used to optimize the path planning strategy under limited parameter updates to ensure that the generated path conforms to human preferences. The DPO mechanism includes: (S5.1) Sparse parameter update strategy: Freeze the GRU and adapter parameters trained in the SFT phase and update them through binary masks , realize only polar coordinate projection matrix Implementation Sparse updates; (S5.2) The preference loss function adopts the DPO loss function: , in is the frozen SFT model, To optimize the strategy, For input, For positive samples, is a negative sample; at the same time, a KL divergence constraint is added to prevent the strategy from deviating too far: 。 7. The method for spatial cognition optimization of large language models with enhanced geometric constraints according to claim 1, characterized in that: The establishment of a multi-dimensional testing framework includes the following: (1) Basic orientation analysis: The geometric mapping accuracy of text instructions is quantified by the polar coordinate error rate model, covering basic capability verification in static scenarios; (2) Dynamic path planning: In a hierarchical maze environment, the first-time planning success rate and path efficiency are calculated to evaluate the model's real-time decision-making ability in a dynamic environment; (3) High-order structural reasoning: Based on the three-dimensional object transformation task of the Mental Rotation Test (MRT), the model's ability to abstractly understand complex spatial relationships is verified through key point matching accuracy and rotation matrix Riemann error.

8. A large language model spatial cognition optimization system enhanced by geometric constraints, characterized by: Used to implement the method according to any one of claims 1 to 7.

9. An electronic device, characterized in that: include: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing computer-executable instructions, wherein the instructions are used to implement the method according to any one of claims 1 to 7 when executed.

Citation Information

Cited By

  • Report generation agent autonomous construction method based on knowledge enhancement large model

    CN122114187A