An environment semantic perception method based on multi-modal collaborative calibration
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NAVAL AVIATION UNIV
- Filing Date
- 2026-05-21
- Publication Date
- 2026-08-07
AI Technical Summary
然而,现有多模态语义感知方法大多通过隐式方式对不同模态信息进行融合,缺乏对多模态特征之间关系的显式建模,容易导致模态间对齐不稳定,从而影响感知结果的准确性
本申请提供了一种基于多模态协同校准的环境语义感知方法,通过获取目标环境区域的视觉图像和语义信息构建环境语义感知模型,环境语义感知模型包括特征提取模块、特征对齐模块、特征校准模块和语义感知预测模块,其中,特征对齐模块通过引入了多模态协同原型实现了对不同模态信息的统一表述和有效融合,提升了环境语义感知模型在复杂场景中的稳定性与感知精度;特征校准模块根据视觉特征,基于多尺度一致性损失,对多模态协同特征进行校准,降低了噪声干扰对语义感知结果的影响,增强了模型对复杂环境变化的适应能力;语义感知预测模块对校准特征进行语义推理,提升了环境语义感知模型的整体感知精度与鲁棒性;利用语义感知损失对环境语义感知模型进行训练,并利用训练好的环境语义感知模型确定目标环境区域中各位置的语义类别,使得本申请在无特定场景先验条件下,有效刻画不同模态之间的协同关系并对其进行校准,提高复杂环境或极端环境的识别精度。
Smart Images

Figure CN122530685A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of multimodal perception technology, and in particular to an environmental semantic perception method based on multimodal collaborative calibration. Background Technology
[0002] Semantic perception tasks aim to achieve semantic-level understanding of targets and their spatial distribution in input images or videos, and are an important foundation for realizing high-level visual understanding and intelligent decision-making. In real-world environments, perception systems often need to simultaneously deal with multiple complex factors, such as changes in ambient lighting, differences in weather conditions, diversity of scene structures, and differences in imaging equipment. These factors can significantly affect the stability and discriminative power of visual features.
[0003] With the development of multimodal technologies, researchers have begun to incorporate auxiliary information such as textual semantics to enhance the understanding capabilities of visual perception models. Multimodal semantic information typically possesses strong abstractness and stability, providing high-level semantic constraints for visual perception. However, most existing multimodal semantic perception methods fuse different modal information implicitly, lacking explicit modeling of the relationships between multimodal features. This can easily lead to unstable alignment between modalities, thus affecting the accuracy of perception results. Furthermore, in complex environments, the contribution of different modal information to the perception task is not consistent, and redundant or noisy information can easily be introduced during the computation process, even interfering with the perception results.
[0004] Therefore, based on the above problems, there is an urgent need to provide an environmental semantic perception method based on multimodal collaborative calibration, which can effectively characterize and calibrate the collaborative relationship between different modalities in multimodal semantic perception, thereby improving the recognition accuracy of complex or extreme environments. Summary of the Invention
[0005] The purpose of this application is to provide an environmental semantic perception method based on multimodal collaborative calibration, which can effectively characterize the collaborative relationship between different modalities in multimodal semantic perception and calibrate them to improve the recognition accuracy of complex or extreme environments.
[0006] To achieve the above objectives, this application provides the following solution: This application provides an environmental semantic awareness method based on multimodal collaborative calibration, including: Obtain visual images and semantic information of the target environment region; the semantic information is category descriptive text. An environmental semantic perception model is constructed based on the visual image and semantic information. The environmental semantic perception model includes a feature extraction module, a feature alignment module, a feature calibration module, and a semantic perception prediction module. The feature extraction module uses a visual encoder and a text encoder to extract features from the visual image and semantic information respectively, obtaining visual features and linguistic features. The feature alignment module aligns the visual features and linguistic features according to a multimodal collaborative prototype, obtaining multimodal collaborative features. The multimodal collaborative prototype is a learnable vector parameter obtained from linguistic feature mapping. The feature calibration module calibrates the multimodal collaborative features based on the visual features and multi-scale consistency loss, obtaining calibrated features. The semantic perception prediction module performs semantic reasoning on the calibrated features to determine the semantic category corresponding to each pixel position in the visual image. The environmental semantic perception model is trained based on semantic perception loss; and the trained environmental semantic perception model is used to determine the semantic category of each location in the target environmental region.
[0007] Optionally, the processing procedure of the feature extraction module includes: Based on the pre-trained visual-language alignment model, a visual encoder and a text encoder are constructed respectively; Visual features are obtained by using a visual encoder to extract features from visual images; A text encoder is used to extract semantic features to obtain linguistic features.
[0008] Optionally, the processing procedure of the feature alignment module includes: Using formula Determine the collaborative weights between visual features and linguistic features; where, For the first Cooperative weights between visual and linguistic features at various scales for function, For the first Visual features at various scales For the first Language representation of alignment at each scale For the first The number of channels for visual features at each scale; By parametrically mapping language features, a multimodal collaborative prototype is obtained; Based on collaborative weights, using the formula Weighted aggregation of multimodal collaborative prototypes yields multimodal collaborative features; among which, For the first Multimodal collaborative features at various scales The number of semantic units in the text. An index for semantic units of text. For the first The feature vectors corresponding to the multimodal collaborative prototypes.
[0009] Optionally, the processing procedure of the feature calibration module includes: Based on similarity measurement, the semantic consistency coefficient between visual features and multimodal collaborative features is determined; The visual features are adaptively reweighted using the semantic consistency coefficient to obtain the calibrated features; Based on multi-scale consistency loss, the alignment relationship between calibration features and multimodal collaborative features is constrained.
[0010] Optionally, determining the semantic consistency coefficient between visual features and multimodal collaborative features based on similarity metrics specifically includes: Using formula Determine the semantic consistency coefficient between visual features and multimodal collaborative features; in, For the first Visual features at various scales and multimodal co-features in spatial location The semantic consistency coefficient at the location, In spatial location First Visual features at various scales In spatial location First Multimodal collaborative features at various scales.
[0011] Optionally, the step of adaptively reweighting the visual features using the semantic consistency coefficient to obtain calibrated features specifically includes: Using formula Obtain calibration characteristics; in, In spatial location First Calibration features at various scales For the first Visual features at various scales and multimodal co-features in spatial location The semantic consistency coefficient at the location, In spatial location First Visual features at various scales.
[0012] Optionally, the constraint on the alignment relationship between calibration features and multimodal collaborative features based on multi-scale consistency loss specifically includes: Using formula Constrain the alignment relationship between the initial calibration features and the multimodal collaborative features; in, For multi-scale consistency loss, For the scale index of the visual encoder, The scale of the visual feature. It is the L2 norm. In spatial location First Calibration features at various scales In spatial location First Multimodal collaborative features at various scales.
[0013] Optionally, training the environment semantic perception model based on semantic perception loss specifically includes: Environmental semantic perception model is used to perform environmental semantic perception and determine the initial prediction results; Based on the initial prediction results and the true semantic categories, the environmental semantic perception model is trained using semantic perception loss to obtain a trained environmental semantic perception model.
[0014] Optionally, the step of using an environmental semantic perception model to perform environmental semantic perception and determine the initial prediction result specifically includes: Using formula Determine the initial prediction results; in, This is the initial prediction result. For environmental semantic perception model, For calibration features, For the scale index of the visual encoder, The scale of the visual feature.
[0015] Optionally, the step of training the environmental semantic perception model based on the initial prediction results and the true semantic categories, using semantic perception loss, to obtain a trained environmental semantic perception model, specifically includes: Using formula Determine semantic perception loss ; in, Let cross-entropy be the loss function. This is the initial prediction result. This represents the true semantic category.
[0016] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides an environmental semantic perception method based on multimodal collaborative calibration. It constructs an environmental semantic perception model by acquiring visual images and semantic information of the target environment region. The environmental semantic perception model includes a feature extraction module, a feature alignment module, a feature calibration module, and a semantic perception prediction module. The feature alignment module introduces a multimodal collaborative prototype to achieve a unified representation and effective fusion of information from different modalities, improving the stability and perception accuracy of the environmental semantic perception model in complex scenes. The feature calibration module calibrates the multimodal collaborative features based on visual features and multi-scale consistency loss, reducing the impact of noise interference on the semantic perception results and enhancing the model's adaptability to complex environmental changes. The semantic perception prediction module performs semantic inference on the calibrated features, improving the overall perception accuracy and robustness of the environmental semantic perception model. The environmental semantic perception model is trained using semantic perception loss, and the trained model is used to determine the semantic categories of each location in the target environment region. This allows the application to effectively characterize and calibrate the collaborative relationships between different modalities without specific scene priors, improving the recognition accuracy in complex or extreme environments. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart illustrating an environmental semantic perception method based on multimodal collaborative calibration in one embodiment of this application. Figure 2 This is a flowchart of an environmental semantic perception method based on multimodal collaborative calibration in one embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] In one exemplary embodiment, such as Figure 1and Figure 2 As shown, an environment semantic awareness method based on multimodal collaborative calibration is provided, including the following S1 to S3. Wherein: S1: Acquire visual images and semantic information of the target environment region.
[0022] Semantic information specifically refers to descriptive text about categories in the environment. The categories are mainly object category names, such as roads, pedestrians, vehicles, and buildings.
[0023] S2: Construct an environmental semantic perception model based on visual images and semantic information.
[0024] The environmental semantic awareness model includes a feature extraction module, a feature alignment module, a feature calibration module, and a semantic awareness prediction module.
[0025] The feature extraction module utilizes a visual encoder and a text encoder to extract features from visual images and semantic information, respectively, to obtain visual features and linguistic features. The processing steps of the feature extraction module include: 1. Based on the pre-trained visual-language alignment model, construct a visual encoder and a text encoder respectively.
[0026] The pre-trained visual-language alignment model can be the CLIP pre-trained model or other pre-trained models with cross-modal alignment capabilities. In this application, the visual branch in the CLIP pre-trained model is used to construct the visual encoder, and the text branch in the CLIP pre-trained model is used to construct the text encoder.
[0027] 2. Use a visual encoder to extract features from the visual image to obtain visual features.
[0028] The visual image (input image) is denoted as ,in, The height of the visual image. The width of the visual image. The number of color channels in a visual image is used as the basis for encoding the image through a visual extraction network, thereby obtaining multi-scale visual features. The visual extraction network employs a pre-trained encoder structure based on visual language alignment (i.e., a visual encoder). In this application, the visual encoder originates from the visual branch of the CLIP pre-trained model, and its parameters are kept frozen during training. The multi-scale visual feature expressions are as follows: ; in, For the first Visual features at various scales For the scale index of the visual encoder, The scale of the visual feature. For the first Visual feature spatial resolution at various scales For the first The number of channels for visual features at each scale.
[0029] 2. Use a text encoder to extract semantic features to obtain language features.
[0030] Semantic information relevant to the semantic perception task is introduced, and this information is encoded using a text encoder to obtain linguistic features. The text encoder, derived from the text branch of the CLIP pre-trained model, maps category names or semantic descriptions to a semantic embedding space aligned with visual features. The expression is as follows: ; in, For the first The language feature vector corresponding to each semantic unit of text. The number of semantic units in the text. An index for semantic units of text. Dimensions of text features.
[0031] This application performs joint feature extraction on the input visual image and semantic information, and uses a visual encoder and a text encoder from the same source to construct a unified visual-language feature representation. This enables visual features and language features to have a good semantic alignment foundation in the initial stage, reduces the difficulty of subsequent multimodal collaborative modeling, and provides basic features for subsequent multimodal information collaboration and calibration.
[0032] The feature alignment module aligns visual and linguistic features based on a multimodal collaborative prototype to obtain multimodal collaborative features. The multimodal collaborative prototype consists of learnable vector parameters obtained from linguistic feature mapping. The processing steps of the feature alignment module include: 1. Determine the collaborative weights between visual features and linguistic features.
[0033] For visual features at any number of scales First, the language features are mapped to the corresponding scale space to obtain scale-aligned language representations. Subsequently, the collaborative weights between visual features and linguistic semantics are calculated using a matching function, as shown in the following formula: ; in, For the first The collaborative weights between visual and linguistic features at each scale are used to characterize the first... At each scale, the degree of correlation between visual and linguistic features reflects the strength of guidance exerted by different semantic units on the current visual features. for function, For the first Visual features at various scales For the first Language representation of alignment at each scale For the first The number of channels for visual features at each scale.
[0034] 2. Perform parameterized mapping on language features to obtain a multimodal collaborative prototype.
[0035] To enhance the stability of multimodal information collaboration, this application introduces a learnable multimodal collaborative prototype for semantic relabeling, thereby avoiding semantic drift in downstream semantic perception tasks by the text encoder. Specifically, the multimodal collaborative prototype is essentially a learnable vector parameter, serving as an intermediate representation between the visual and linguistic modalities. Obtained from linguistic feature mapping, it aligns visual and linguistic features generated by the CLIP pre-trained models (visual encoder and text encoder), and performs task-related semantic adaptation based on this. The expression is as follows: ; in, It serves as a multimodal collaborative prototype index, and also as an index for textual semantic units. This refers to the number of multimodal co-prototypes, which is also the number of text semantic units; that is, the number of multimodal co-prototypes is consistent with the number of text semantic units. For the first The feature vectors corresponding to the multimodal collaborative prototypes.
[0036] In this application, visual features at any scale are... Through multimodal collaborative prototyping Attention-matching operations between visual features yield preliminary visual-guided feature representations, which are then used to update or enhance the visual semantic components contained in the multimodal collaborative prototype. Through this interactive process, the multimodal collaborative prototype not only incorporates prior linguistic information but also adaptively absorbs visual feature information from the current input sample.
[0037] 3. Based on the collaborative weight, the multimodal collaborative prototypes are weighted and aggregated to obtain multimodal collaborative features.
[0038] The formula for calculating multimodal collaborative features is as follows: ; in, For the first Multimodal collaborative features at various scales The number of semantic units in the text. An index for semantic units of text. For the first The feature vectors corresponding to the multimodal collaborative prototypes.
[0039] This application calculates the collaborative weights between visual and linguistic features through a matching mechanism and introduces a learnable multimodal collaborative prototype set to explicitly characterize the semantic correspondence between different modalities and further strengthen their consistency in the semantic space. Finally, it generates multimodal collaborative features, which enables this application to inherit the cross-modal alignment capability of the CLIP pre-trained model while achieving adaptive modeling of task-related semantic information, providing a more stable feature representation for subsequent multi-scale information calibration and semantic perception prediction.
[0040] The feature calibration module is used to calibrate multimodal collaborative features based on visual features and multi-scale consistency loss to obtain calibrated features.
[0041] To further improve the semantic consistency and stability of feature representations at different scales, this application performs multi-scale information calibration processing on multi-scale visual features and multi-modal collaborative features, thereby reducing the impact of scale variations and local noise on semantic perception results. The feature calibration module's processing includes: 1. Based on similarity measurement, determine the semantic consistency coefficient between visual features and multimodal collaborative features.
[0042] A scale-consistency mapping is constructed based on similarity metrics, which is applicable to arbitrary scales. The formula for calculating the semantic consistency coefficient between visual features and multimodal collaborative features is as follows: ; in, For the first Visual features at various scales and multimodal co-features in spatial location The semantic consistency coefficient at the location, In spatial location First Visual features at various scales In spatial location First Multimodal collaborative features at various scales.
[0043] In this application, the semantic consistency coefficient ranges from [value range missing]. Or located after normalization The larger the value of the interval, the more the visual features at that location conform to the stable semantic structure guided by multimodal collaboration at the semantic level.
[0044] 2. Adaptively reweight the visual features using the semantic consistency coefficient to obtain calibrated features.
[0045] Specifically, the formula for calculating the calibration feature is as follows: ; in, In spatial location First Calibration features at various scales For the first Visual features at various scales and multimodal co-features in spatial location The semantic consistency coefficient at the location, In spatial location First Visual features at various scales.
[0046] By reweighting, feature regions with higher semantic consistency are given greater contribution weight in subsequent perception processes, while regions with lower semantic consistency are relatively suppressed.
[0047] 3. Based on multi-scale consistency loss, the alignment relationship between calibration features and multi-modal collaborative features is constrained.
[0048] Specifically, to maintain overall semantic consistency at the scale level, this application introduces a multi-scale consistency constraint loss to constrain the overall semantic alignment relationship between calibration features and multimodal collaborative features at different scales. Its formal definition is: ; in, For multi-scale consistency loss, For the scale index of the visual encoder, The scale of the visual feature. It is the L2 norm. In spatial location First Calibration features at various scales In spatial location First Multimodal collaborative features at various scales.
[0049] Multi-scale consistency constraint loss is used to guide features at different scales to maintain consistency at the global semantic level, thereby avoiding semantic drift between multi-scale features. Through the above similarity-based consistency calibration mechanism, this application can achieve semantic alignment and stable modeling of multi-scale features without explicitly introducing feature differencing or error mapping operations.
[0050] The semantic-aware prediction module is used to perform semantic reasoning on calibration features to determine the semantic category corresponding to each pixel position in the visual image.
[0051] S3: Train the environmental semantic perception model based on semantic perception loss; and use the trained environmental semantic perception model to determine the semantic category of each location in the target environmental region.
[0052] S3 specifically includes: S31: Use the environmental semantic perception model to perform environmental semantic perception and determine the initial prediction results.
[0053] Let the environmental semantic perception model be The initial prediction result obtained by the environmental semantic perception model is: ; in, This is the initial prediction result, specifically the semantic category prediction result corresponding to each pixel position in the visual image. For environmental semantic perception model, For calibration features.
[0054] S32: Based on the initial prediction results and the true semantic category, train the environmental semantic perception model using semantic perception loss to obtain the trained environmental semantic perception model.
[0055] The environmental semantic perception model is trained end-to-end to obtain a well-trained environmental semantic perception model with stable semantic expression capabilities. During the training process, supervised learning is used to compare the initial prediction results with the real semantic labels / real semantic categories. Apply constraints and construct a semantic-aware prediction loss. The calculation formula is as follows: ; in, The cross-entropy loss function can be used, but other supervised losses suitable for semantic awareness tasks can also be used depending on the specific circumstances.
[0056] This application utilizes multi-scale consistency constraint loss and semantic perception prediction loss to jointly optimize the environmental semantic perception model, constructing an overall training objective function. The calculation formula is as follows: ; in, The weighting coefficient is used to balance the contribution weights of different loss terms. The weighting coefficient can be set according to different task requirements or training stages.
[0057] S33: Use the trained environmental semantic perception model to determine the semantic category of each location in the target environment region.
[0058] A well-trained environmental semantic perception model can gradually learn semantically consistent and stable feature representations through the combined effect of multimodal information collaboration and multi-scale information calibration, thereby achieving stable and accurate semantic perception of complex scenes and obtaining the semantic categories of each location in the target environment region.
[0059] This application constructs an initially aligned multimodal feature space through a vision-language pre-trained model. Combined with a learnable multimodal collaborative prototype, it achieves refined collaborative modeling of visual and linguistic features. Through multi-scale information calibration and consistency constraint mechanisms, it enhances the semantic stability and perception accuracy of the environmental semantic perception model under different scales and complex scene conditions. Furthermore, through a joint training optimization strategy, it improves the overall perception accuracy and robustness of the environmental semantic perception model. This application does not rely on specific scene priors and is applicable to various complex visual semantic perception tasks, enabling its widespread application in autonomous driving, intelligent security, robot perception, and unmanned system navigation.
[0060] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0061] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An environmental semantic perception method based on multimodal collaborative calibration, characterized in that, The environment semantic perception method based on multimodal collaborative calibration includes: Obtain visual images and semantic information of the target environment region; the semantic information is category descriptive text. An environmental semantic perception model is constructed based on the visual image and semantic information. The environmental semantic perception model includes a feature extraction module, a feature alignment module, a feature calibration module, and a semantic perception prediction module. The feature extraction module uses a visual encoder and a text encoder to extract features from the visual image and semantic information respectively, obtaining visual features and linguistic features. The feature alignment module aligns the visual features and linguistic features according to a multimodal collaborative prototype, obtaining multimodal collaborative features. The multimodal collaborative prototype is a learnable vector parameter obtained from linguistic feature mapping. The feature calibration module calibrates the multimodal collaborative features based on the visual features and multi-scale consistency loss, obtaining calibrated features. The semantic perception prediction module performs semantic reasoning on the calibrated features to determine the semantic category corresponding to each pixel position in the visual image. The environmental semantic perception model is trained based on semantic perception loss; and the trained environmental semantic perception model is used to determine the semantic category of each location in the target environmental region.
2. The environmental semantic perception method based on multimodal collaborative calibration according to claim 1, characterized in that, The processing steps of the feature extraction module include: Based on the pre-trained visual-language alignment model, a visual encoder and a text encoder are constructed respectively; Visual features are obtained by using a visual encoder to extract features from visual images; A text encoder is used to extract semantic features to obtain linguistic features.
3. The environmental semantic perception method based on multimodal collaborative calibration according to claim 1, characterized in that, The processing steps of the feature alignment module include: Using formula Determine the collaborative weights between visual features and linguistic features; where, For the first Cooperative weights between visual and linguistic features at various scales for function, For the first Visual features at various scales For the first Language representation of alignment at each scale For the first The number of channels for visual features at each scale; By parametrically mapping language features, a multimodal collaborative prototype is obtained; Based on collaborative weights, using the formula Weighted aggregation of multimodal collaborative prototypes yields multimodal collaborative features; among which, For the first Multimodal collaborative features at various scales The number of semantic units in the text. An index for semantic units of text. For the first The feature vectors corresponding to the multimodal collaborative prototypes.
4. The environmental semantic perception method based on multimodal collaborative calibration according to claim 1, characterized in that, The processing procedure of the feature calibration module includes: Based on similarity measurement, the semantic consistency coefficient between visual features and multimodal collaborative features is determined; The visual features are adaptively reweighted using the semantic consistency coefficient to obtain the calibrated features; Based on multi-scale consistency loss, the alignment relationship between calibration features and multimodal collaborative features is constrained.
5. The environmental semantic perception method based on multimodal collaborative calibration according to claim 4, characterized in that, The determination of the semantic consistency coefficient between visual features and multimodal collaborative features based on similarity measurement specifically includes: Using formula Determine the semantic consistency coefficient between visual features and multimodal collaborative features; in, For the first Visual features at various scales and multimodal co-features in spatial location The semantic consistency coefficient at the location, In spatial location First Visual features at various scales In spatial location First Multimodal collaborative features at various scales.
6. The environmental semantic perception method based on multimodal collaborative calibration according to claim 4, characterized in that, The step of adaptively reweighting visual features using semantic consistency coefficients to obtain calibrated features specifically includes: Using formula Obtain calibration characteristics; in, In spatial location First Calibration features at various scales For the first Visual features at various scales and multimodal co-features in spatial location The semantic consistency coefficient at the location, In spatial location First Visual features at various scales.
7. The environmental semantic perception method based on multimodal collaborative calibration according to claim 4, characterized in that, The constraint on the alignment relationship between calibration features and multimodal collaborative features based on multi-scale consistency loss specifically includes: Using formula Constrain the alignment relationship between the initial calibration features and the multimodal collaborative features; in, For multi-scale consistency loss, For the scale index of the visual encoder, The scale of the visual feature. It is the L2 norm. In spatial location First Calibration features at various scales In spatial location First Multimodal collaborative features at various scales.
8. The environmental semantic perception method based on multimodal collaborative calibration according to claim 1, characterized in that, The training of the environment semantic perception model based on semantic perception loss specifically includes: Environmental semantic perception model is used to perform environmental semantic perception and determine the initial prediction results; Based on the initial prediction results and the true semantic categories, the environmental semantic perception model is trained using semantic perception loss to obtain a trained environmental semantic perception model.
9. The environmental semantic perception method based on multimodal collaborative calibration according to claim 8, characterized in that, The step of using an environmental semantic perception model to perform environmental semantic perception and determine the initial prediction result specifically includes: Using formula Determine the initial prediction results; in, This is the initial prediction result. For environmental semantic perception model, For calibration features, For the scale index of the visual encoder, The scale of the visual feature.
10. The environmental semantic perception method based on multimodal collaborative calibration according to claim 8, characterized in that, The step of training the environmental semantic perception model based on the initial prediction results and the true semantic categories, using semantic perception loss, to obtain a trained environmental semantic perception model, specifically includes: Using formula Determine semantic perception loss ; in, Let cross-entropy be the loss function. This is the initial prediction result. This represents the true semantic category.