Elevator map generation method, system and equipment based on semantic recognition and image segmentation large model and medium

By using a large model based on semantic recognition and image segmentation, elevator status documents are automatically generated, solving the problems of manual dependence and semantic bias in traditional elevator inspection. This enables accurate positioning and structured description of elevator components, improving inspection efficiency and database utilization.

CN120894636AActive Publication Date: 2025-11-04SICHUAN SPECIAL EQUIP INSPECTION & RES INST

Patent Information

Application Number
CN202511068668.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-11-04
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

In current elevator inspection, traditional visual data post-processing relies on manual intervention, resulting in low inspection timeliness, cross-personal visual semantic understanding bias, low database utilization, and an inability to accurately query the status of elevator components.

Method used

A method based on semantic recognition and image segmentation large model is adopted. The segmentation mask map of elevator components is generated by pre-trained semantic segmentation model. Combined with cross-modal fusion features and text generation model, a structured elevator status document is generated. The text description is optimized by multi-source calibration mechanism.

Benefits of technology

It achieves automated conversion of elevator images into standardized documents, accurately locates key elevator components, generates machine-readable structured documents, solves the problems of subjective bias in manual interpretation and unstructured data retrieval, and improves detection efficiency and database utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894636A_ABST
    Figure CN120894636A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of elevator image processing, in particular to an elevator image generation method, system and device based on a semantic recognition and image segmentation large model and a medium, and the method comprises the steps: obtaining an original image of an elevator detection scene and carrying out the preprocessing; the preprocessed original image is processed based on a pre-trained semantic segmentation model, and a segmentation mask graph of the elevator component is generated; respectively inputting the original image and the segmented mask image into an image encoder for feature extraction, and fusing the extracted features to generate cross-modal fusion features; based on the cross-modal fusion features, elevator part associated information is analyzed through a text generation model, and initial text description is generated; and optimizing the initial text description by utilizing a multi-source calibration mechanism, and generating a standardized elevator state document containing structured metadata. The objective of the invention is to realize automatic semantic analysis and structured report generation of elevator images.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of elevator image processing technology, in particular to an elevator image text generation method, system, device and medium based on semantic recognition and large image segmentation model. BACKGROUND

[0002] As a key link in the safety supervision system of special equipment, the real-time monitoring and whole life cycle management of the elevator are directly related to public safety. In order to meet the standards of traceability and objectivity in the supervision requirements, the industry generally uses intelligent terminals or special imaging equipment to visually collect key detection components, so as to carry information such as elevator component state and detection operation track in an intuitive visual form, which constitutes an important basis for equipment safety evaluation.

[0003] In the existing elevator detection scene, the post-processing of traditional visual data still relies on manual guidance. The semantic gap between unstructured visual data and structured reports leads to the need for detection personnel to extract component state features through frame-level analysis and manually transcribe them into standardized descriptions. This process significantly reduces the detection timeliness due to redundant operations. Secondly, the deviation in visual semantic understanding across personnel makes it difficult to consistently describe key abnormal states such as door gap exceeding the standard and local deformation of guide rails. Finally, the disconnection between low-level visual features and high-level semantics results in a database that only supports basic retrieval based on timestamps and equipment IDs, making it impossible to accurately query based on component types and fault features, leading to low database utilization. SUMMARY

[0004] To achieve automatic semantic analysis and structured report generation of elevator images, the present application provides an elevator image text generation method, system, device and medium based on semantic recognition and image segmentation large model, the technical solutions adopted are as follows: The technical solution of the first aspect of the present application provides an elevator image text generation method based on semantic recognition and image segmentation large model, the method comprises: Obtain the original image of the elevator detection scene and perform preprocessing; Process the preprocessed original image based on a pre-trained semantic segmentation model to generate a segmentation mask image of the elevator components; Input the original image and the segmentation mask image into an image encoder respectively for feature extraction, and fuse the extracted features to generate cross-modal fusion features; Based on the cross-modal fusion features, use a text generation model to analyze elevator component related information to generate an initial text description; Optimize the initial text description using a multi-source calibration mechanism to generate a standardized elevator state document containing structured metadata.

[0005] Further, the preprocessed original image is processed based on a pre-trained semantic segmentation model to generate a segmentation mask image of the elevator component, including: A standard data set of multi-class segmentation of elevator components is constructed, and a semantic segmentation model is trained based on the standard data set; The preprocessed original image is multi-level down-sampled by an encoder to extract high-level semantic features; The high-level semantic features are up-sampled by a decoder and spliced with the features of the corresponding layer of the encoder; Based on the class probability distribution of the output layer, a multi-class mask image is generated by threshold determination.

[0006] Further, the original image and the segmentation mask image are respectively input into an image encoder for feature extraction, and the extracted features are fused to generate cross-modal fusion features, including: A first encoder is used to extract a global environment feature vector from the original image; A second encoder is used to extract an elevator component spatial positioning feature vector from the segmentation mask image; The global environment feature vector and the elevator component spatial positioning feature vector are added to generate cross-modal fusion features.

[0007] Further, based on the cross-modal fusion features, a text generation model is used to analyze elevator component associated information to generate an initial text description, including: Based on the cross-modal fusion features, a spatial topology vector is analyzed by the self-attention mechanism of the Transformer; According to the spatial topology vector and the elevator component knowledge base, a text sequence containing component name, position description and abnormal label is generated.

[0008] Further, a multi-source calibration mechanism is used to optimize the initial text description to generate a standardized elevator state document containing structured metadata, including: The segmentation mask image is converted into a component position text sequence; A pre-defined elevator component label is introduced as a calibration reference; The edit distance difference between the component position text sequence and the pre-defined elevator component label is calculated; The semantic similarity between the component position text sequence and the initial text description is calculated; The component position description is jointly corrected according to the edit distance difference and the semantic similarity.

[0009] Further, the expression of the multi-source calibration mechanism loss function is:

[0010] In the formula, represents the closed-loop calibration loss; component position text sequence semantic alignment weight coefficient with initial text description component position text sequence semantic difference loss with initial text description predefined elevator component label semantic difference loss with initial text description component position text sequence structural alignment weight coefficient with predefined elevator component label component position text sequence text structure difference loss with predefined elevator component label

[0011] Further, the standardized elevator state document comprises: component state description partitioned by car system, door system and traction system; metadata binding elevator registration code, detection timestamp and shaft position coordinates; stored in JSON format and establishing component type index field.

[0012] The technical solution of the second aspect of the application provides an elevator graph text generation system based on a large model of semantic recognition and image segmentation, which adopts the elevator graph text generation method based on the large model of semantic recognition and image segmentation of the first aspect of the application. The system comprises: An image acquisition module configured to acquire and preprocess original images of an elevator detection scene; A segmentation module configured to process the preprocessed original images based on a pre-trained semantic segmentation model to generate a segmentation mask image of elevator components; A feature fusion module configured to input the original images and the segmentation mask image into an image encoder respectively for feature extraction, and fuse the extracted features to generate cross-modal fusion features; A text generation module configured to parse elevator component associated information based on the cross-modal fusion features to generate an initial text description using a text generation model; A calibration output module configured to optimize the initial text description using a multi-source calibration mechanism to generate a standardized elevator state document containing structured metadata.

[0013] ​​​​​The technical scheme of the third aspect of the present application provides an electronic device, which comprises a processor and a memory connected with the processor; wherein the memory stores instructions executable by the processor, and the instructions are executed by the processor to enable the processor to execute the steps of the elevator diagram text generation method based on semantic recognition and image segmentation large model according to the first aspect of the present application.

[0014] The technical scheme of the fourth aspect of the present application provides a computer readable storage medium, which stores a program for implementing the elevator diagram text generation method based on semantic recognition and image segmentation large model, and the program is executed by a processor to implement the steps of the elevator diagram text generation method based on semantic recognition and image segmentation large model according to the first aspect of the present application.

[0015] The present application has the following advantages: The elevator diagram text generation method based on semantic recognition and image segmentation large model provided by the present application realizes the automatic conversion of elevator images to standardized documents in the elevator state document automatic generation scene; the pre-trained semantic segmentation model is used to realize the accurate positioning of the key components of the elevator, the cross-modal fusion features are combined with the Transformer architecture to analyze the component topology relationship and generate standardized descriptions; the segmentation mask, the predefined label and the generated text are optimized by the multi-source calibration mechanism, the subjective bias of manual interpretation is solved, the machine interpretable structured document with semantic indexing capability is generated, and the problem of non-structured visual data that cannot be deeply searched is avoided. BRIEF DESCRIPTION OF DRAWINGS

[0016] In order to more clearly illustrate the technical schemes and advantages of the embodiments of the present application or the prior art, the drawings needed in the following embodiment or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.

[0017] Figure 1 The method flowchart of the elevator diagram text generation method, system, device and medium based on semantic recognition and image segmentation large model provided by an embodiment of the present application; Figure 2 The model architecture schematic diagram of the elevator diagram text generation method based on semantic recognition and image segmentation large model provided by an embodiment of the present application; Figure 3 The structure schematic diagram of the elevator diagram text generation method, system, device and medium based on semantic recognition and image segmentation large model provided by an embodiment of the present application. Detailed Implementation

[0018] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of an elevator image-to-text method, system, device, and medium based on a large model of semantic recognition and image segmentation proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0020] The following description, in conjunction with the accompanying drawings, details a specific scheme for an elevator image-to-text method, system, device, and medium based on a large model of semantic recognition and image segmentation provided by this invention.

[0021] Please see Figure 1 This document illustrates a flowchart of an elevator image-to-text method, system, device, and medium based on a large-scale model of semantic recognition and image segmentation, according to an embodiment of the present invention. The method includes: Step S100: Acquire and preprocess the original images of the elevator inspection scenario. Specifically, in the elevator status inspection scenario, using image acquisition equipment, such as a fixed camera inside the elevator or a handheld shooting device, high-definition images of key areas such as the elevator car interior, shaft, and machine room are captured according to predetermined inspection time nodes. The acquired images are transmitted to the system server via wired or wireless network. After receiving the images, the server first performs a unified conversion of the image format, transforming various original image formats, such as JPEG and PNG, into a system-compatible standard format. The same preprocessing operations as the U-Net model training data are then performed, including: resizing: using bicubic interpolation, the image and corresponding mask are uniformly scaled to 512×512 pixels; for images whose aspect ratio is inconsistent with the target size, zeros are padded at the edges to ensure that the image does not deform during scaling while maintaining the integrity of the content; normalization: the image pixel values ​​are normalized and mapped to the [0,1] interval to eliminate the differences in brightness and contrast between different images, enabling the model to learn image features more stably and accelerate the convergence speed during training and inference.

[0022] Step S200: Process the pre-processed original image based on the pre-trained semantic segmentation model to generate a segmentation mask map of the elevator components; Step S200 specifically includes: Step S210: Construct an elevator component multi-class segmentation standard dataset, and train a semantic segmentation model based on the standard dataset; specifically, images and corresponding masks of different models of elevators are obtained, and key components such as a car and a guide rail are labeled; then the dataset is divided into a training set, a validation set, and a test set at a ratio of 8:1:1. The images are preprocessed, scaled to 512x512 pixels, normalized to the [0, 1] interval, and enhanced in data diversity through rotation and flipping; see Figure 2 In this embodiment, a U-Net architecture is adopted as the segmentation model, which includes an encoder and a decoder. The encoder extracts semantic features through convolution and pooling. The decoder restores the spatial resolution through deconvolution and connects with the corresponding layer features of the encoder. The output layer uses a 1x1 convolution to generate a component class probability map. Step S220: The preprocessed original image is subjected to multi-level down-sampling through the encoder to extract high-level semantic features; specifically, the preprocessed 512x512x3 image is input into the encoder of the pre-trained U-Net, and feature extraction is completed through four down-sampling modules in sequence: each module first extracts features from the image through two 3x3 convolution layers, the feature maps after convolution are normalized through a BatchNorm layer to reduce internal covariate shift, and then a ReLU activation function is introduced to introduce nonlinearity and enhance the expression ability of the model. Subsequently, a 2x2 max-pooling layer with a step size of 2 reduces the size of the feature map by half and doubles the number of channels, gradually compressing the spatial information of the image and extracting high-level semantic features, including global classes and spatial correlations of elevator components. Step S230: The high-level semantic features are up-sampled through the decoder and spliced with the features of the corresponding layers of the encoder; specifically, the 32x32x512 high-level features output by the encoder are input into the decoder, and the spatial resolution is gradually restored through four up-sampling modules: The first up-sampling module: the input features are up-sampled through a 2x2 deconvolution layer with 256 convolution kernel numbers and a step size of 2, and a 64x64x256 feature map is output; the 64x64x256 feature map output by the third down-sampling module of the encoder is spliced in the channel dimension through a skip connection to fuse high-level semantics and middle-level detail features; the features are refined through two 3x3 convolution layers, and a 64x64x256 feature map is output. The second to fourth up-sampling modules: the structure is consistent with that of the first up-sampling module, the number of deconvolution kernels is halved in sequence, and the sizes after up-sampling are 128x128x128, 256x256x64, and 512x512x32, respectively, which are spliced and fused with the features of the second, first, and zero layers of the encoder, i.e., the features of the original image after initial convolution. The bottom-level details such as edges and textures preserved by the encoder are transmitted through a skip connection, combined with the semantic information of the decoder, so that the model can accurately locate the component outline while restoring the spatial resolution.

[0023] Step S240: based on the output layer category probability distribution, a multi-category mask graph is generated by threshold decision; specifically, the decoder finally outputs a 512x512x32 feature graph, which is mapped to the number of categories by a 1x1 convolution layer to obtain a 512x512x32 category score graph; the Softmax activation function is applied to the score graph to calculate the probability of each pixel belonging to each category, and a 512x512x32 probability distribution graph is output; the probability threshold of this embodiment is set to 0.5, and the category with the highest probability is selected for each pixel, and if the probability is greater than or equal to 0.5, it is determined to be the category, otherwise it is determined to be the background, and finally a 512x512 multi-category mask graph is generated, marking the outlines and positions of each elevator component.

[0024] This embodiment realizes accurate segmentation of elevator components by constructing a specialized dataset and optimizing the U-Net model architecture, and the output multi-category mask graph not only preserves the complete spatial outline of the components, but also accurately marks the category information, providing a reliable component positioning basis for subsequent cross-modal feature fusion. At the same time, through the quality control mechanism, the stability of the segmentation result is ensured, effectively solving the problems of component recognition confusion, blurred edge positioning in complex scenes, etc., and significantly improving the bottom feature analysis accuracy of the graph generation text system.

[0025] Step S300: input the original image and the segmentation mask graph into the image encoder respectively for feature extraction, and fuse the extracted features to generate cross-modal fusion features; Step S300 specifically includes: Step S310: use the first encoder to extract a global environment feature vector from the original image; the coding module uses a pre-trained ResNet-50 as the backbone network to extract features from the image: after the original image is input into ResNet-50, it successively passes through the convolution layer, the pooling layer and the residual block in the network; by continuously convolving and pooling, the local features, global features and semantic information at different levels of the image are gradually extracted, and finally a high-dimensional feature vector is output, which contains the overall structure and semantic information of the original image, such as the internal layout of the car, the lighting conditions of the shaft, the overall distribution of the machine room equipment, etc. Step S320: use the second encoder to extract an elevator component spatial positioning feature vector from the segmentation mask graph; similarly, a pre-trained ResNet-50 is used as the second encoder to extract features from the 512x512 segmentation mask graph generated in step S200, which is a single channel, and the pixel value represents the component category: The mask image is first converted into a 512x512x3 three-channel image through dimension expansion, the single-channel value is copied to three channels, the input format of the ResNet-50 is adapted, and the second encoder is input. After the same network structure processing as step S310: 7x7 convolution, 3x3 pooling and four residual block groups in turn, the final 1x1x2048 feature vector is output through global average pooling. The vector focuses on the spatial distribution features of the elevator components, such as "the car door is located on the left side of the image" and "the traction machine is located in the upper center of the image", and makes up for the lack of component positioning details in the original image features. Since the segmentation mask image mainly reflects the position and contour information of the components, the feature vector obtained after encoding highlights the spatial distribution features of the components. The feature vectors obtained by encoding the original image and the segmentation mask image are merged in the subsequent steps to provide a more comprehensive image semantic representation for text generation; Step S330: adding the global environment feature vector and the spatial positioning feature vector of the elevator components to generate a cross-modal fusion feature; using feature addition fusion: the feature vector obtained by encoding the mask image and the feature vector obtained by encoding the original image are added in the corresponding dimensions to realize the fusion of the two features. This step integrates the fine local features from the target components of the mask image and the global environment features from the original image to generate a more comprehensive and representative fusion feature vector, which provides rich information for the subsequent Q-Former to analyze the relationship between components.

[0026] This embodiment uses a dual-channel encoding and feature fusion strategy to construct a cross-modal feature representation that takes into account both global and local features: the first encoder extracts global environment features from the original image, providing scene context, and the second encoder extracts spatial positioning features from the mask image, strengthening component details. The addition of the two features realizes complementary enhancement of semantic information and avoids the one-sidedness of single-image features or mask features. The fusion feature contains both the category and position information of the elevator components and is associated with the environment in which they are located, laying a data foundation for the subsequent Q-Former model to accurately analyze image semantics and generate text descriptions that conform to the actual scene, effectively solving the technical limitations of traditional single-modal features that cannot simultaneously consider "what component" and "where is it located", addressing the one-sidedness of single-modal feature information, and significantly improving the semantic analysis capabilities of the image-to-text system for complex elevator scenes.

[0027] Step S400: based on the cross-modal fusion feature, using a text generation model to analyze elevator component-related information to generate an initial text description; Step S400 specifically includes: Step S410: Based on the cross-modal fusion feature, the spatial topology vector is analyzed by the self-attention mechanism of the Transformer; the fused feature vector is used as the input of the Q-Former model; the Q-Former model is based on the Transformer architecture, and uses the self-attention mechanism to perform deep analysis on the input image feature: the fused feature vector is converted into the input sequence of the Q-Former through the linear projection layer, and the sequence length is set according to the feature dimension; at the same time, a set of learnable query vectors are initialized for the model, which are used to focus on the key component features. The attention weight between the query vector and the input sequence is calculated through the self-attention mechanism, and the attention weight matrix reflects the correlation strength of different component features, for example, the spatial proximity between the car door and the guide rail will correspond to a higher weight. After processing by the multi-layer Transformer encoder, the encoder includes a self-attention layer and a feedforward neural network, and outputs a topology vector that encodes the spatial relationship between components. This vector integrates global environment and local component information in the fused feature, and provides a quantitative basis for semantic association analysis.

[0028] Step S420: According to the spatial topology vector and the elevator component knowledge base, a text sequence containing component names, position descriptions and abnormality labels is generated; the elevator component knowledge base is constructed: 32 types of component information are stored according to the safety specifications for elevator manufacturing and installation, including standard names, typical position descriptions, abnormal state characteristics and associated components; the knowledge base uses a triple structure for storage; based on the spatial topology vector and the elevator component knowledge base, a structured text is generated through a fully connected layer and an LLM Decoder, and the specific process is as follows: the fully connected layer receives the topology vector output by the Q-Former, adjusts the feature dimension through nonlinear transformation, so that it is adapted to the feature space of the large language model, ensures that the image feature and the language feature distribution of the LLM are consistent, and improves the feature analysis efficiency. The adapted feature is used as the conditional input of the LLM, combined with the pre-trained elevator component knowledge base, which contains structured data such as component name attributes and typical position descriptions. The LLM (Large Language Model) uses a self-recurrent generation strategy to predict the next token word by word from the starting label: The model first analyzes the component category from the topology vector and matches the standard name in the knowledge base; Generate position description based on spatial relationship information; If an abnormal feature is detected from the abnormal region label of the segmentation mask in step S200, an abnormality label is added; The final generated initial text sequence includes component name shaft coordinates and abnormal state; component name such as "car door", "drive motor", position such as "car top", "shaft side wall" and other information description text. The generation process adopts the autoregressive method, starts from the starting mark, and gradually predicts the next word until the complete text sequence is generated. LLM is based on the massive language knowledge and text generation ability learned by pre-training, combined with input features and label guide information, to generate the final output text. In the example, the output is consistent with the input label "A drive motor", and in actual application, more detailed and accurate text description will be generated according to the image features and labels, such as motor model, appearance state, working mode association description, etc. This embodiment realizes accurate mapping from image features to structured text through the spatial relationship analysis of Q-Former and the semantic generation ability of LLM. The self-attention mechanism effectively captures the topological association between components, solving the problem that a single feature is difficult to express spatial relationship. The feature adaptation of the fully connected layer ensures the effective transmission of cross-modal information, and the text generated by LLM in combination with the professional knowledge base not only contains standardized component names and position information, but also accurately describes abnormal states, laying a semantic foundation for subsequent document standardization and avoiding the subjectivity and non-standardization of manual description.

[0029] Step S500: optimizing the initial text description by using a multi-source calibration mechanism to generate a standardized elevator state document containing structured metadata; wherein the standardized elevator state document includes: component state description partitioned according to car system, door system, and drive system; metadata binding elevator registration code, detection timestamp, and shaft position coordinates; stored in JSON format and establishing a component type index field; Step S500 specifically includes: Step S510: converting the segmentation mask map into a component position text sequence; based on the segmentation mask map generated in step S200, the text sequence is generated through the following process: Region analysis: for each connected region in the mask map, corresponding to a single elevator component, extract its minimum bounding rectangle coordinates and convert them into shaft relative coordinates, such as "1.5m from the bottom of the shaft, 0.3m to the left".

[0030] Text mapping: combined with the preset component category and position description template, such as "[component name] is located at [shaft position], coordinate range", convert each region into structured text segment. For example, the "car door" region in the mask map corresponds to the text "car door is located at the front side of the shaft, coordinate range (50, 200) - (300, 400)"; Sequence integration: sort according to component system, car system, door system, etc. to generate a complete component position text sequence , containing all detected components and their position information; Step S520: Introduce predefined elevator component labels as calibration reference; Specifically, predefined elevator component labels Based on elevator industry standard construction, including: component name labels: such as "car door", "hoist machine", "guide rail", etc., one-to-one corresponding to the categories in the segmentation mask image; position description labels: such as "left side of the hoistway", "central top of the car", "below the hoist machine", etc., standardizing position expression; relationship description labels: such as "parallel to XX component", "above XX component", etc., defining the standard spatial relationship between components; these labels are stored in the system label library as a reference for subsequent text calibration.

[0031] Step S530: Calculate the edit distance difference between the component position text sequence and the predefined elevator component label; Edit distance is used to quantify the structural difference between , the calculation process is as follows: for each text segment in , the corresponding label in , calculate the minimum single-character editing operation required to convert to ; for example , "car door" needs to be replaced with "car door" in , the edit distance is 1; "lock device" is missing and needs to be inserted, the edit distance increases by 1; the structural alignment loss of component segmentation text can be expressed as:

[0032] In the formula, the smaller the value, the higher the structural consistency between and ; Step S540: Calculate the semantic similarity between the component position text sequence and the initial text description; use cosine similarity to measure the semantic association with the initial text description generated in step S400, the process is as follows: use the Sentence-BERT model to encode and respectively to obtain semantic vectors and , with a dimension of 768, where is the Sentence-BERT encoding of , is the Sentence-BERT encoding of ; Step S550: jointly correct the component position description according to the edit distance difference and the semantic similarity; the expression of the multi-source calibration mechanism loss function is as follows:

[0033] In the formula, represents the closed-loop calibration loss; represents the semantic alignment weight coefficient of the component position text sequence and the initial text description , which is 0.6, and the semantic consistency of the two is preferentially ensured; represents the semantic difference loss of the component position text sequence and the initial text description , which is calculated in the same way as , that is, , is the F1 score of a text pair calculated based on a pre-trained BERT model, and the smaller the value is, the more consistent the semantics is; represents the semantic difference loss of the predefined elevator component label and the initial text description ; represents the structure alignment weight coefficient of the component position text sequence and the predefined elevator component label , which is 0.4; represents the text structure difference loss of the component position text sequence and the predefined elevator component label , which directly uses the edit distance value, and the smaller the value is, the more matched the structure is; wherein, the first term : the weight is 0.6, which strengthens the semantic consistency of and , for example, to correct the description conflict of “car door position”; the second term is used to ensure that meets the industry standard label terms, for example, to correct “door machine” to “landing door driving device”. The third term is used to constrain the structure of to be consistent with , for example, to supplement the missing “safety hook” description; The embodiment implements text correction based on the step-by-step loss function and the total loss function, and the total loss function is as follows:

[0034] In the formula, represents the total loss function; represents the semantic alignment loss of the label and the initial text; represents the semantic synergy loss of and , that is wherein and are Sentence-BERT encoding vectors of and respectively; the Adam optimizer is adopted to minimize , and the parameters of the text generation model, i.e., Q-Former and LLM, are updated through backpropagation; wherein the component position description joint correction process comprises: if the deviation positioning is : it is indicated that and the semantic deviation is significant, and the conflicting fragments are located by comparing the semantic vectors and for example describes “the traction machine is on the left side” while describes “on the right side”, and the spatial positioning feature of is taken as a reference for correction; if : it is indicated that there is a non-standard term, and is referred to for replacement, for example, “ladder box” is uniformly corrected to “car”; if : it is prompted that there is a structural deficiency, and the component description is supplemented according to , for example, the related text of “speed limiter” is added; wherein the text structured and corrected text is sorted according to “car system-door system-traction system-guide system-safety protection system”, and is secondarily sorted according to the importance of components in the same system. The embedded metadata includes the elevator registration code, the detection timestamp and the shaft position coordinates; finally, it is stored in the JSON format, and the component type index field is established, for example, “car_system” corresponds to the car system text, and “door_system” corresponds to the door system text, which supports quick retrieval according to the component type and the system category.

[0035] In this embodiment, the multi-source calibration mechanism integrates structural constraints and semantic constraints, and the initial text description is corrected based on the joint optimization of the edit distance and the semantic similarity, so that the generated standardized elevator state document not only conforms to the industry standard term specification but also accurately reflects the spatial relationship of components. Meanwhile, the structured metadata design realizes efficient retrieval and management, effectively solves the problem of strong subjectivity and format disorder of manual reports, and provides a standardized and traceable text basis for the whole life cycle safety management of elevators.

[0036] In summary, the elevator graph text generation method based on semantic recognition and image segmentation large model provided in the application is based on the elevator component document automatic generation technology of U-Net image recognition and Q-Former text generation, and a closed-loop optimization system from image segmentation to text generation is constructed by introducing a multi-dimensional loss function. First, the U-Net model is input after the elevator image is preprocessed to realize accurate segmentation of the elevator components. Then, the original image and the segmented mask image are encoded by using ResNet-50, the encoded results are fused, the fused feature vectors are input into the Q-Former model to generate description text, the text generated by the Q-Former and the text generated by the U-Net are optimized and fused to generate an optimized output, and a standardized document is output. The U-Net model is optimized in terms of component segmentation accuracy by using a combination of generalized Dice Loss and Cross Entropy Loss, the semantic alignment of the text generated by the Q-Former and the input label is strengthened by using BERTScore Loss and contrastive learning loss, and the collaborative optimization of the U-Net segmentation result, the input label and the text generated by the Q-Former is realized by using a triangular closed-loop loss.

[0037] In practical applications, compared with the traditional manual review and induction method, the method can greatly improve the efficiency of elevator state document generation, and the time from processing a single image to generating a document is shortened from an average of 10 minutes to a few seconds. Through the collaborative optimization of U-Net and Q-Former, the matching accuracy of the text description and the actual elevator component state is improved. At the same time, the closed-loop optimization system constructed by the multi-dimensional loss function gives the system self-learning and self-adaptive ability, reduces the manual participation link, and effectively improves the automation and standardization level of elevator state inspection and document management. Please refer to Figure 3 which shows a structure schematic diagram of an elevator graph text generation system based on a semantic recognition and image segmentation large model provided by an embodiment of the application, the system comprising: An image acquisition module configured to acquire and preprocess an original image of an elevator detection scene; A segmentation module configured to process the preprocessed original image based on a pre-trained semantic segmentation model to generate a segmentation mask image of elevator components; A feature fusion module configured to input the original image and the segmentation mask image into an image encoder respectively for feature extraction, and fuse the extracted features to generate cross-modal fusion features; A text generation module configured to generate initial text description based on the cross-modal fusion features by using a text generation model to analyze elevator component associated information; A calibration output module configured to optimize the initial text description by using a multi-source calibration mechanism to generate a standardized elevator state document containing structured metadata.

[0038] The technical solution of the third aspect of the present application provides an electronic device, comprising: a processor and a memory connected in communication with the processor; wherein the memory stores instructions executable by the processor, and the instructions are executed by the processor to enable the processor to perform the steps of the elevator diagram text generation method based on semantic recognition and image segmentation large model according to the technical solution of the first aspect of the present application.

[0039] The technical solution of the fourth aspect of the present application provides a computer readable storage medium, and the computer readable storage medium stores a program for implementing an elevator diagram text generation method based on semantic recognition and image segmentation large model. The program for implementing the elevator diagram text generation method based on semantic recognition and image segmentation large model is executed by a processor to implement the steps of the elevator diagram text generation method based on semantic recognition and image segmentation large model according to the technical solution of the first aspect of the present application.

[0040] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description, and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are possible or can be advantageous.

[0041] Each embodiment in the specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment focuses on the difference from other embodiments.

Claims

1. An elevator image-to-text method based on a large-scale semantic recognition and image segmentation model, characterized in that, The method includes: Acquire the original image of the elevator inspection scene and perform preprocessing; The pre-processed original image is processed based on a pre-trained semantic segmentation model to generate a segmentation mask image of the elevator components. The original image and the segmentation mask image are respectively input into the image encoder for feature extraction, and the extracted features are fused to generate cross-modal fusion features; Based on the cross-modal fusion features, an initial text description is generated by parsing the elevator component association information using a text generation model. The initial text description is optimized using a multi-source calibration mechanism to generate a standardized elevator status document containing structured metadata.

2. The elevator drawing-to-text method as described in claim 1, characterized in that, The pre-trained semantic segmentation model processes the pre-processed original image to generate a segmentation mask map of the elevator components, including: Construct a multi-category segmentation standard dataset for elevator components, and train a semantic segmentation model based on the standard dataset; The encoder performs multi-level downsampling on the preprocessed original image to extract high-level semantic features; The high-level semantic features are upsampled by the decoder and then concatenated with the features of the corresponding layer of the encoder. Based on the category probability distribution of the output layer, a multi-class mask image is generated by threshold determination.

3. The elevator drawing-to-text method as described in claim 1, characterized in that, The original image and the segmentation mask image are respectively input into an image encoder for feature extraction, and the extracted features are fused to generate cross-modal fusion features, including: The first encoder is used to extract global environmental feature vectors from the original image; The second encoder is used to extract the spatial localization feature vectors of elevator components from the segmentation mask image; The global environment feature vector is added to the spatial positioning feature vector of the elevator components to generate cross-modal fusion features.

4. The elevator drawing-to-text method as described in claim 3, characterized in that, Based on the aforementioned cross-modal fusion features, an initial text description is generated by parsing the elevator component association information using a text generation model, including: Based on cross-modal fusion features, spatial topological vectors are parsed using the self-attention mechanism of Transformer; Based on spatial topology vectors and elevator component knowledge base, a text sequence containing component names, location descriptions, and anomaly markers is generated.

5. The elevator drawing-to-text method according to any one of claims 1 to 4, characterized in that, The initial text description is optimized using a multi-source calibration mechanism to generate a standardized elevator status document containing structured metadata, including: Convert the segmentation mask image into a sequence of component location text; Introduce predefined elevator component labels as calibration benchmarks; Calculate the difference in edit distance between the component location text sequence and the predefined elevator component labels; Calculate the semantic similarity between the component location text sequence and the initial text description; The component location description is corrected based on the difference in edit distance and semantic similarity.

6. The elevator drawing-to-text method as described in claim 5, characterized in that, The expression for the loss function of the multi-source calibration mechanism is: In the formula, This indicates the closed-loop calibration loss; Text sequence indicating component location Compared with the initial text description Semantic alignment weight coefficients; Text sequence indicating component location Compared with the initial text description Semantic difference loss; Indicates predefined elevator component labels Compared with the initial text description Semantic difference loss; Text sequence indicating component location With predefined elevator component labels The structural alignment weight coefficient; Text sequence indicating component location With predefined elevator component labels Text structure difference loss.

7. The elevator drawing-to-text method as described in claim 5, characterized in that, The standardized elevator status document includes: Component status descriptions by car system, door system, and traction system; Metadata including elevator registration code, detection timestamp, and shaft location coordinates; Store and create a component type index field in JSON format.

8. An elevator image-to-text system based on a large-scale semantic recognition and image segmentation model, characterized in that: The elevator image-to-text method based on a large model of semantic recognition and image segmentation as described in any one of claims 1 to 7, the system comprising: The image acquisition module is configured to acquire and preprocess the original images of the elevator detection scene. The segmentation module is configured to process the preprocessed original image based on a pre-trained semantic segmentation model to generate a segmentation mask image of the elevator components. The feature fusion module is configured to input the original image and the segmentation mask image into the image encoder for feature extraction, and then fuse the extracted features to generate cross-modal fusion features. The text generation module is configured to generate an initial text description by parsing the elevator component association information based on the cross-modal fusion features and using a text generation model. The calibration output module is configured to optimize the initial text description using a multi-source calibration mechanism to generate a standardized elevator status document containing structured metadata.

9. An electronic device, characterized in that, The electronic device includes: a processor and a memory communicatively connected to the processor; wherein the memory stores instructions executable by the processor, the instructions being executed by the processor to enable the processor to perform the steps of the elevator image-to-text method based on a large model of semantic recognition and image segmentation as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program for implementing an elevator image-to-text method based on a large model of semantic recognition and image segmentation. The program for implementing the elevator image-to-text method based on a large model of semantic recognition and image segmentation is executed by a processor to implement the steps of the elevator image-to-text method based on a large model of semantic recognition and image segmentation as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Medical image report generation method based on hidden space image-text matching

    CN118609745A

  • Robust multi-modal image segmentation method and system based on instance perception query

    CN119672342A

  • Engineering cost information optimization integration method based on Internet big data service

    CN120216749A

  • System and method for detecting elevator maintenance behavior in elevator hoistway

    US20200130999A1

Cited By

  • Environment document processing method and device based on large model, equipment and storage medium

    CN121094084A

  • Elevator security behavior detection method, device and equipment and computer storage medium

    CN121637347A

  • Laser alignment guide rail straightness dynamic compensation measurement method based on multi-sensor fusion

    CN122258796A

  • A dynamic compensation measurement method for straightness of a laser collimation guide rail

    CN122258796B