Field maintenance method and device, electronic equipment and computer program product

By employing multimodal data processing methods, the problem of insufficient automation and intelligence in resource management during on-site maintenance of communication networks has been solved, enabling more efficient resource identification, comparison, and retrieval, and improving the level of automation and intelligence in on-site maintenance.

CN121967255APending Publication Date: 2026-05-01CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER
Filing Date
2024-10-29
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

The existing communication network field maintenance has problems with insufficient automation and intelligence in equipment identification and resource management, relying on manual methods for comparison, verification and validation.

Method used

A multimodal data processing method is adopted to acquire multimodal data in field maintenance. A multimodal sub-model for the target task type is determined by a multimodal scenario selection model. Resource identification, comparison or retrieval of multimodal data is performed. Data alignment, fusion and cross-validation are performed using the multimodal sub-model. Decisions are made in conjunction with a decision model.

Benefits of technology

It improves the ability to identify, compare, and retrieve resources in on-site maintenance of communication networks, enhances the level of automation and intelligence, and achieves more efficient resource management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121967255A_ABST
    Figure CN121967255A_ABST
Patent Text Reader

Abstract

The invention relates to a field maintenance method and device, electronic equipment and a computer program product, and belongs to the technical field of artificial intelligence. The method comprises the following steps: acquiring multi-modal data in field maintenance; according to a multi-modal data processing method, carrying out resource identification or resource comparison or resource retrieval on the multi-modal data to obtain a processing result of the multi-modal data; and performing field maintenance according to the processing result of the multi-modal data. According to the invention, resource identification, resource comparison or resource retrieval processing is carried out on the multi-modal data in field maintenance through the multi-modal data processing method, so that the automation and intelligence level of communication network field maintenance can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and more specifically, to a field maintenance method, a field maintenance device, an electronic device, and a computer program product. Background Technology

[0002] On-site maintenance of communication networks requires frequent identification, verification, configuration management, and other tasks related to the identification, verification, configuration, and management of resource information and status of equipment, boards, fiber optic cables, and the equipment room and on-site environment.

[0003] Currently, AI-based image recognition methods can be used in some scenarios to assist in the identification of devices, etc. However, this method is limited by factors such as the quality of the photo images and the accuracy of the training data. There are still many situations where manual comparison, verification, and validation are required, which cannot meet the needs of automated and intelligent on-site maintenance of communication networks.

[0004] Therefore, there is an urgent need in this field for a field maintenance method that can improve the automation and intelligence level of field maintenance of communication networks.

[0005] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0006] The purpose of this disclosure is to provide a field maintenance method, field maintenance device, electronic device, and computer program product, thereby improving the automation and intelligence level of field maintenance of communication networks to at least a certain extent.

[0007] According to a first aspect of this disclosure, a field maintenance method is provided, comprising:

[0008] Acquire multimodal data during on-site maintenance;

[0009] According to the multimodal data processing method, resource identification, resource comparison, or resource retrieval are performed on the multimodal data to obtain the processing result of the multimodal data;

[0010] On-site maintenance is performed based on the processing results of the multimodal data.

[0011] In one exemplary embodiment of this disclosure, the multimodal data includes at least two of the following: image data, video data, text data, and structured data.

[0012] In one exemplary embodiment of this disclosure, the step of performing resource identification, resource comparison, or resource retrieval on the multimodal data according to the multimodal data processing method to obtain the processing result of the multimodal data includes:

[0013] The target task is input into the multimodal scene selection model. Based on the output of the multimodal scene selection model, a multimodal sub-model of the target type corresponding to the target task is determined from multiple types of multimodal sub-models. The types of multimodal sub-models include multimodal sub-models for resource identification, multimodal sub-models for resource comparison, and multimodal sub-models for resource retrieval.

[0014] The multimodal data is processed according to the multimodal sub-model of the target type to obtain the processing result of the multimodal data, which includes resource identification result, resource comparison result, or resource retrieval result.

[0015] In one exemplary embodiment of this disclosure, when the target task is a resource identification task, the step of processing the multimodal data according to the multimodal sub-model of the target type to obtain the processing result of the multimodal data includes:

[0016] The multimodal data is aligned and / or fused according to the multimodal sub-model for resource identification to obtain the resource identification result, which includes the classification and identification results of the resource type, resource occupancy and resource status of the multimodal data.

[0017] In one exemplary embodiment of this disclosure, when the target task is a resource comparison task, the step of processing the multimodal data according to the multimodal sub-model of the target type to obtain the processing result of the multimodal data includes:

[0018] The multimodal data is cross-validated according to the multimodal sub-model used for resource comparison to obtain the resource comparison result. The resource comparison result includes the judgment result of whether the multimodal data matches, and the matched multimodal data.

[0019] In one exemplary embodiment of this disclosure, when the target task is a resource retrieval task, the step of processing the multimodal data according to the multimodal sub-model of the target type to obtain the processing result of the multimodal data includes:

[0020] The multimodal data is subjected to cross-modal retrieval based on the multimodal sub-model used for resource retrieval to obtain the resource retrieval results, which include the cross-modal retrieval results of the multimodal data.

[0021] In one exemplary embodiment of this disclosure, the step of performing data cross-validation on the multimodal data according to the multimodal sub-model for resource comparison to obtain the resource comparison result includes:

[0022] Based on the multimodal sub-model for resource comparison, cross-validation is performed on multiple pairs of interrelated multimodal data to obtain the associated resource comparison results.

[0023] In one exemplary embodiment of this disclosure, the step of performing data cross-validation on the multimodal data according to the multimodal sub-model for resource comparison to obtain the resource comparison result includes:

[0024] Compare the current multimodal data with historical multimodal data;

[0025] If the current multimodal data is identical to the historical multimodal data, then the resource comparison result indicating no difference is directly returned.

[0026] If the current multimodal data differs from the historical multimodal data, then the multimodal data is cross-validated according to the multimodal sub-model used for resource comparison to obtain the resource comparison result.

[0027] In one exemplary embodiment of this disclosure, the step of performing cross-modal retrieval on the multimodal data according to the multimodal sub-model for resource retrieval to obtain the resource retrieval result includes:

[0028] Based on the multimodal sub-model for resource retrieval, cross-modal retrieval is performed on multiple pairs of interrelated multimodal data to obtain the associated resource retrieval results.

[0029] In one exemplary embodiment of this disclosure, processing the multimodal data according to the multimodal sub-model of the target type to obtain the processing result of the multimodal data includes:

[0030] The multimodal data is input into the multimodal sub-model of the target type to obtain the sub-model output result of the multimodal sub-model of the target type;

[0031] The output of the sub-model is input into the decision model, and the processing result of the multimodal data is obtained according to the decision strategy of the decision model.

[0032] According to a second aspect of this disclosure, a field maintenance device is provided, comprising:

[0033] The data acquisition module is used to acquire multimodal data during on-site maintenance.

[0034] The data processing module is used to perform resource identification, resource comparison, or resource retrieval on the multimodal data according to the multimodal data processing method, and obtain the processing result of the multimodal data;

[0035] The field maintenance module is used to perform field maintenance based on the processing results of the multimodal data.

[0036] According to a third aspect of this disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to perform the field maintenance method described in any of the preceding claims by executing the executable instructions.

[0037] According to a fourth aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the field maintenance method described in any of the preceding claims.

[0038] The exemplary embodiments disclosed herein can have the following beneficial effects:

[0039] In the field maintenance method of this exemplary embodiment, multimodal data in the field maintenance is acquired, and resource identification, comparison, or retrieval is performed on the multimodal data according to a multimodal data processing method to obtain the processing result of the multimodal data. Then, field maintenance is performed based on the processing result of the multimodal data. The field maintenance method in this exemplary embodiment, by performing resource identification, comparison, or retrieval on the multimodal data in the field maintenance through a multimodal data processing method, can improve the resource identification, comparison, and retrieval capabilities of communication network field maintenance, as well as the automation and intelligence level of communication network field maintenance.

[0040] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0041] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0042] Figure 1 A flowchart illustrating an exemplary embodiment of the field maintenance method of this disclosure is shown;

[0043] Figure 2 A flowchart illustrating a field maintenance method based on a multimodal approach according to a specific embodiment of the present disclosure is shown.

[0044] Figure 3 A flowchart illustrating a multimodal data processing method according to an exemplary embodiment of this disclosure is shown.

[0045] Figure 4 The schematic diagram illustrates a framework of a multimodal sub-model for resource identification according to a specific embodiment of the present disclosure;

[0046] Figure 5 A block diagram of a field maintenance apparatus according to an exemplary embodiment of the present disclosure is shown;

[0047] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing the embodiments of the present disclosure is shown. Detailed Implementation

[0048] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this disclosure more comprehensive and complete, and to fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., can be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0049] Furthermore, the accompanying drawings are merely illustrative of this disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0050] This example implementation first provides a field maintenance method. (See reference...) Figure 1 As shown, the above-mentioned on-site maintenance method may include the following steps:

[0051] Step S110. Obtain multimodal data during on-site maintenance.

[0052] Step S120. According to the multimodal data processing method, perform resource identification, resource comparison, or resource retrieval on the multimodal data to obtain the processing result of the multimodal data.

[0053] Step S130. Perform on-site maintenance based on the processing results of the multimodal data.

[0054] In the field maintenance method of this exemplary embodiment, multimodal data in the field maintenance is acquired, and resource identification, comparison, or retrieval is performed on the multimodal data according to a multimodal data processing method to obtain the processing result of the multimodal data. Then, field maintenance is performed based on the processing result of the multimodal data. The field maintenance method in this exemplary embodiment, by performing resource identification, comparison, or retrieval on the multimodal data in the field maintenance through a multimodal data processing method, can improve the resource identification, comparison, and retrieval capabilities of communication network field maintenance, as well as the automation and intelligence level of communication network field maintenance.

[0055] Below, in conjunction with Figures 2 to 4 The steps described above in this example implementation will be explained in more detail.

[0056] In step S110, multimodal data from on-site maintenance is acquired.

[0057] In this example implementation, field maintenance refers to the process of maintaining equipment in a communication network field. The multimodal data in field maintenance can include at least two of the following: image data, video data, text data, and structured data. Specifically, the multimodal data in communication network field maintenance may include, but is not limited to: image and video data from visible light cameras, image and video data from infrared cameras, image and video data from depth cameras, handwritten and printed text on equipment panels and labels used in the field maintenance, and structured and unstructured data such as resource data, alarm data, and log data from external related systems used in the field maintenance.

[0058] In step S120, according to the multimodal data processing method, resource identification, resource comparison, or resource retrieval are performed on the multimodal data to obtain the processing result of the multimodal data.

[0059] In this example implementation, a multimodal model can be used to identify, compare, and retrieve resources from multimodal data during on-site maintenance. The multimodal model includes a scenario selection model, a multimodal sub-model, and a decision model.

[0060] Figure 2 A flowchart illustrating a field maintenance method based on a multimodal approach according to a specific embodiment of this disclosure is shown. Figure 2 As shown, multimodal data can include data of various modalities such as images, videos, text, and structured data. The specific sources of multimodal data are as follows:

[0061] Resource information for on-site maintenance corresponding to images and videos from visible light cameras, infrared cameras, and depth cameras, including but not limited to: rack model, equipment model, rack location and availability, rack internal slot availability, equipment location, equipment internal slot availability, equipment port availability, equipment indicator light color and switch status, equipment button switch status, fiber optic or network cable connection routing, etc.

[0062] The printed and handwritten text information on the equipment panel and labels corresponds to the on-site maintenance resource information, including but not limited to: printed text near the boards, switches, ports, indicator lights, etc. on the equipment panel, and images and text information on the labels on the fiber optic cables and network cables;

[0063] Resource information, alarm information, and log information in the relevant management system for on-site maintenance of communication networks, including but not limited to: resource information such as device model, device slot occupancy status, board model, device port occupancy status, etc.; alarm information such as alarm name, alarm time, and the corresponding device board port information, etc.; and log information such as log name, log type, and the corresponding device board port information related to device configuration, etc.

[0064] In such Figure 2 The framework shown includes, in addition to the multimodal model, corresponding downstream target tasks and matched multimodal data. The downstream target tasks can include multimodal classification and recognition tasks for resources, multimodal comparison tasks for resource data, and cross-modal retrieval tasks for resource data.

[0065] In this example implementation, as Figure 3 As shown, according to the multimodal data processing method, resource identification, resource comparison, or resource retrieval are performed on the multimodal data to obtain the processing results. Specifically, this may include the following steps:

[0066] Step S310. Input the target task into the multimodal scene selection model, and determine the target type multimodal sub-model corresponding to the target task from multiple types of multimodal sub-models based on the output of the multimodal scene selection model.

[0067] The types of multimodal sub-models include multimodal sub-models for resource identification, multimodal sub-models for resource comparison, and multimodal sub-models for resource retrieval.

[0068] In this example implementation, the specific type of multimodal sub-model can be selected using a scenario selection model.

[0069] The selection method for the scenario selection model can be based on external task input, such as based on the type of downstream task; or it can be based on self-supervised learning, which learns from the input data to determine which type of multimodal model should be used for the current scenario and then calls that multimodal model for processing.

[0070] Step S320. Process the multimodal data according to the multimodal sub-model of the target type to obtain the processing result of the multimodal data.

[0071] The processing results of multimodal data include resource identification results, resource comparison results, or resource retrieval results.

[0072] In this example implementation, when the target task is resource identification, multimodal data can be aligned and / or fused according to the multimodal sub-model used for resource identification to obtain resource identification results. The resource identification results include the classification and identification results of resource type, resource occupancy and resource status of the multimodal data.

[0073] Figure 4 The diagram illustrates a framework of a multimodal sub-model for resource identification according to a specific embodiment of the present disclosure.

[0074] The multimodal sub-model for resource identification can provide alignment and fusion of multiple modal data to achieve classification and identification of multimodal data, and is mainly used in downstream task scenarios such as resource identification.

[0075] Among them, two or more modal data from the communication network field maintenance can be selected for multimodal alignment and fusion.

[0076] This can be achieved using conventional multimodal deep learning methods, including but not limited to: multimodal fusion methods such as early fusion, late fusion, and hybrid fusion; attention mechanisms and other methods can be used to enhance multimodal alignment and fusion capabilities; for example, for unimodal representation methods, images can use ViT (Vision Transformer), text can use BERT (Bidirectional Encoder Representations from Transformers), and videos can use time-series frames. Multimodal alignment and fusion methods, such as CLIP (Contrastive Language-Image Pre-Training), ALBEF (Align Before Fuse), and VL-BERT (Vision-Language-BERT), are used for image-text multimodal fusion. This example implementation does not limit the specific image, video, and text alignment and fusion method used in this model.

[0077] Among them, different implementation instances of multimodal sub-models for resource identification can be provided to correspond to the different modal data and different fusion methods mentioned above.

[0078] In this example implementation, by comprehensively utilizing data from multiple modalities in the field maintenance of the communication network, such as images or videos from visible light cameras, infrared cameras, and depth cameras, as well as handwritten or printed text on equipment tags, and structured and unstructured data such as resource data, alarm data, and log data from relevant management systems, this multimodal data is aligned and fused to improve the ability to identify resources in the field maintenance of the communication network, including the ability to automatically and intelligently identify, compare, and verify resource types, resource occupancy, and resource status.

[0079] In this example implementation, when the target task is a resource comparison task, the multimodal data can be cross-validated according to the multimodal sub-model used for resource comparison to obtain the resource comparison result. The resource comparison result includes the judgment result of whether the multimodal data matches, as well as the matched multimodal data.

[0080] The multimodal sub-model for resource comparison can provide cross-validation of multiple modal data to achieve multimodal data comparison, and is mainly used in downstream task scenarios such as resource comparison.

[0081] This can be achieved through deep learning-based image-text comparison and matching methods to verify the comparison of multiple modalities of data, such as ITC (Image-text contrastive) and ITM (Image-text matching). Generative contrastive judgment methods, such as GAN (Generative Adversarial Networks), can also be used for the recognition and verification of multiple modalities of data. Other methods may also be employed. Corresponding field maintenance scenarios include, for example, comparing equipment photos (images) with resource system resource data (text), or comparing equipment port photos (images) with cable labels (printed text), and so on.

[0082] In this example implementation, data cross-validation can be performed on multiple pairs of interrelated multimodal data based on the multimodal sub-model used for resource comparison to obtain the associated resource comparison results.

[0083] For specific scenarios involving on-site maintenance of communication networks, the multimodal sub-model for resource comparison can also support the correlation matching and verification function of multiple image-text pairs. For example, for multiple correlated data in the structured data of external systems such as resource systems, alarm systems, and log systems, such as device model, device port type, device port occupancy status, cable label connected to the device port, occupancy status of the peer device port connected to the other end of the cable, peer device port type, peer device model, etc., the corresponding multimodal data such as images and text should be able to be matched based on the above correlation information.

[0084] In this example implementation, the current multimodal data can also be compared with historical multimodal data. If there is no difference between the current multimodal data and the historical multimodal data, the resource comparison result with no difference is directly fed back. If there is a difference between the current multimodal data and the historical multimodal data, the multimodal data is cross-validated according to the multimodal sub-model used for resource comparison to obtain the resource comparison result.

[0085] For certain special scenarios in the on-site maintenance of communication networks, simplified multimodal models can also be adopted. For example, a simplified judgment method can be used, which involves text-image association and comparison of the current image with historical images. An example scenario is as follows: In a scenario where resource data comparison is performed on the port occupancy of a device in a computer room, the historical image corresponding to the device to be compared is found through the text description of the comparison task. Then, through the computer room monitoring video, monitoring videos from the time the historical image was captured are found. Image frames are extracted at intervals and compared with the historical images. If there is no difference, a no-difference report is given. If there is a difference, the image with the difference is subsequently compared again by calling the scenario selection model in the decision model to select other multimodal sub-models.

[0086] Among them, different implementation instances of multimodal sub-models for resource comparison can be provided to correspond to the different cross-validation methods mentioned above.

[0087] In this example implementation, by performing multimodal alignment and fusion, conversion and comparison of multiple modal data such as images, videos, texts and structured data in the field maintenance of the communication network, the resource comparison capability in the field maintenance of the communication network is improved, including the ability to compare resource types, resource occupancy and resource status.

[0088] In this example implementation, when the target task is a resource retrieval task, cross-modal retrieval of multimodal data can be performed based on the multimodal sub-model used for resource retrieval to obtain resource retrieval results, which include cross-modal retrieval results of multimodal data.

[0089] The multimodal sub-model for resource retrieval can provide cross-modal retrieval capabilities for various modal data, and is mainly used in downstream task scenarios such as resource retrieval.

[0090] Multimodal data that has undergone multimodal model alignment, fusion, and cross-validation, such as image-text matching data and image-text-structured data matching data, can serve as a source of multimodal matching data to support cross-modal retrieval. Through multimodal sub-models used for resource retrieval, cross-modal retrieval can be provided, including text-based image retrieval, image-based text retrieval, and retrieval of corresponding text and images through structured data. Corresponding downstream task scenarios include: in on-site maintenance, providing matching device port images based on text descriptions of device port information in work orders; providing matching images of device port cable connections based on cable information from device ports; and providing text descriptions of device panel information based on on-site photographs, etc., providing reference for on-site maintenance.

[0091] Matching data from multiple modalities after multimodal alignment and fusion and multimodal cross-validation can serve as a data source for cross-modal retrieval, as one of the input information sources for scene selection models, as a data source for training and updating multimodal models, and so on.

[0092] In this example implementation, cross-modal retrieval can be performed on multiple pairs of interrelated multimodal data based on the multimodal sub-model used for resource retrieval, and related resource retrieval results can be obtained.

[0093] For specific scenarios involving on-site maintenance of communication networks, the multimodal sub-model for resource retrieval can also support multimodal association retrieval of multiple image-text pairs. For example, for multiple associated data in structured data from external systems such as resource systems, alarm systems, and log systems, such as equipment model, equipment port type, equipment port occupancy status, cable labels connected to the equipment port, occupancy status of the peer device port connected to the other end of the cable, peer device port type, and peer device model, the corresponding multimodal data such as images and text should be able to be retrieved based on the aforementioned association information.

[0094] Among them, different implementation instances of multimodal sub-models for resource retrieval can be provided to correspond to the different retrieval methods mentioned above.

[0095] In this example implementation, cross-modal data retrieval capabilities are provided based on multimodal data that has been compared through methods such as multimodal alignment and fusion, such as image-text matching data. This enhances resource retrieval capabilities in on-site maintenance, such as retrieving on-site equipment images through text or retrieving text through images, thereby effectively improving the automation and intelligence level of on-site maintenance work.

[0096] In this example implementation, when obtaining the processing result of multimodal data, the multimodal data can also be input into the multimodal sub-model of the target type to obtain the sub-model output result of the multimodal sub-model of the target type, and then the sub-model output result can be input into the decision model to obtain the processing result of multimodal data according to the decision strategy of the decision model.

[0097] If multiple multimodal sub-models are used in multimodal data processing, a decision model can be used to make decisions based on the outputs of each sub-model to obtain the final processing result. The selection of the decision model can be achieved by using external decision inputs to formulate the decision strategy, or by using self-supervised learning to perform self-supervised learning based on the output data of the multimodal model to provide a decision strategy.

[0098] In this example implementation, multiple rounds of judgment can also be performed, from the decision model to the scenario selection model, to improve the accuracy of the judgment.

[0099] In step S130, on-site maintenance is performed based on the processing results of the multimodal data.

[0100] The final output of multimodal data processing can include the output results of multimodal classification and identification tasks for resources, the output results of multimodal comparison tasks for resource data, and the output results of cross-modal retrieval tasks for resource data.

[0101] The output of the multimodal classification and identification task for resources includes the identified resource types, resource occupancy, and resource status for on-site maintenance, such as: equipment model, the availability of internal slots, the availability of equipment ports, the color and on / off status of equipment indicator lights, the identification of optical fibers and network cables, and the connection direction of optical fibers or network cables.

[0102] The output of the multimodal comparison task for resource data includes whether the comparison results are matches, non-matches, and matched multimodal data. Examples include image-text matching data, image-text-structured data matching data, and so on. Specific examples corresponding to on-site maintenance scenarios include matching data between equipment photos (images) and resource data (text) from the resource system, matching data between equipment port photos (images) and cable labels (printed text), and, for example, the associated matching data of multiple image-text pairs: image-text pairs of local equipment model, image-text pairs of local equipment port type, image-text pairs of equipment port occupancy status, image-text pairs of cable labels connected to the equipment port, image-text pairs of occupancy status of the other end of the cable connected to the other end of the equipment port, image-text pairs of the other end equipment port type, and image-text pairs of the other end equipment model.

[0103] The output of the cross-modal retrieval task for resource data includes: cross-modal retrieval results. For example, in the case of on-site maintenance of communication networks, based on the text description of equipment port information in the work order, a matching equipment port image is retrieved; based on the cable information of the equipment port, an image showing the cable connection status of the matching equipment port is retrieved; and based on the equipment information captured on-site, a text description of the equipment panel information is provided. Another example is the multimodal association retrieval of multiple image-text pairs. For multiple associated image-text pairs such as: image-text pair of local equipment model, image-text pair of local equipment port type, image-text pair of equipment port occupancy status, image-text pair of cable label connected to the equipment port, image-text pair of the occupancy status of the other end of the cable connected to the other end equipment port, image-text pair of the other end equipment port type, and image-text pair of the other end equipment model, cross-modal association retrieval can be performed based on the above association information.

[0104] In this example implementation, a multimodal approach is used to align, fuse, and verify data from multiple modalities in field maintenance. This enables the identification, comparison, and verification of resource types, resource occupancy, and resource status during field maintenance, thereby improving the automation and intelligence capabilities of field maintenance of communication networks.

[0105] This example embodiment also provides a field maintenance system based on a multimodal method. This system can include the aforementioned multimodal model and can be deployed within relevant field maintenance systems. For example, the system can input visible light, infrared, and depth images and videos from cameras such as infrared cameras and depth cameras, as well as printed and handwritten text. It can also obtain corresponding resource, alarm, and log information from relevant systems and input it into the multimodal model. Furthermore, the system analyzes and outputs the corresponding multimodal classification and recognition task results for resources, the multimodal comparison task results for resource data, and the cross-modal retrieval task results for resource data from the multimodal model.

[0106] In one specific embodiment, multimodal alignment, fusion, and cross-validation can be performed on the image and video of an indicator light on the device panel, the text description near the indicator light, resource information and alarm information in related systems.

[0107] For example, the location of an indicator light on the device panel can be used to identify a specific function of a port corresponding to that indicator light; the text description near the indicator light can also be used to identify a specific function of a port corresponding to that indicator light; the on / off state of the indicator light can be used to identify the corresponding alarm information; and the resource information and alarm information of the relevant system can be used to identify the alarm information corresponding to that indicator light. The above information is then aligned, fused, and verified using a multimodal approach. The steps are as follows:

[0108] (1) The camera captures an image or video of an indicator light on the device panel, including: the position of the indicator light, its on / off state, etc.;

[0109] (2) Obtain the printed or handwritten text near the panel indicator light, such as the port name number corresponding to the indicator light;

[0110] (3) Obtain from the resources, alarms, and logs of the relevant system the text or structured data of the resource information related to the port of the indicator light, and the structured or unstructured data of the alarm information corresponding to the on / off state of the indicator light;

[0111] (4) Input the images and videos in (1), the printed and handwritten texts in (2), and the structured or unstructured data in (3) into the multimodal model, and analyze and output the corresponding resource information from the multimodal model.

[0112] (5) The output analysis results can be provided to relevant systems for on-site maintenance, thereby assisting in the resource management of on-site maintenance.

[0113] In this example implementation, resource types, resource occupancy, and resource status can be identified based on on-site maintenance images or videos using a multimodal approach. Multimodal verification methods are employed to identify, compare, and verify resource types, resource occupancy, and resource status. For example, multimodal verification can be performed using images or videos from visible light cameras, infrared cameras, depth cameras, images of handwritten or printed text on equipment tags, and resource data, alarm data, or log data from relevant external systems, thereby providing resource identification, resource comparison, and resource retrieval capabilities.

[0114] It should be noted that although the steps of the method in this disclosure are described in a specific order in the accompanying drawings, this does not require or imply that the steps must be performed in that specific order, or that all the steps shown must be performed to achieve the desired result. Additional or alternative steps may be omitted, multiple steps may be combined into one step, and / or a step may be broken down into multiple steps.

[0115] Furthermore, this disclosure also provides a field maintenance device. (See reference) Figure 5 As shown, the field maintenance device may include a data acquisition module 510, a data processing module 520, and a field maintenance module 530. Wherein:

[0116] The data acquisition module 510 can be used to acquire multimodal data during on-site maintenance;

[0117] The data processing module 520 can be used to perform resource identification, resource comparison, or resource retrieval on multimodal data according to the multimodal data processing method, and obtain the processing results of multimodal data;

[0118] The field maintenance module 530 can be used to perform field maintenance based on the processing results of multimodal data.

[0119] In some exemplary embodiments of this disclosure, multimodal data includes at least two of the following: image data, video data, text data, and structured data.

[0120] In some exemplary embodiments of this disclosure, the data processing module 520 may include a sub-model determination unit and a sub-model processing unit. Wherein:

[0121] The sub-model determination unit can be used to input the target task into the multimodal scene selection model, and determine the target type multimodal sub-model corresponding to the target task from multiple types of multimodal sub-models based on the output of the multimodal scene selection model. The types of multimodal sub-models include multimodal sub-models for resource identification, multimodal sub-models for resource comparison, and multimodal sub-models for resource retrieval.

[0122] The sub-model processing unit can be used to process multimodal data according to the multimodal sub-model of the target type to obtain the processing results of the multimodal data, including resource identification results, resource comparison results, or resource retrieval results.

[0123] In some exemplary embodiments of this disclosure, the sub-model processing unit may include a sub-model identification unit, which can be used to perform data alignment and / or data fusion on multimodal data according to the multimodal sub-model used for resource identification, to obtain resource identification results. The resource identification results include classification and identification results of resource type, resource occupancy and resource status of multimodal data.

[0124] In some exemplary embodiments of this disclosure, the sub-model processing unit may further include a sub-model comparison unit, which can be used to perform data cross-validation on multimodal data based on the multimodal sub-model used for resource comparison, and obtain resource comparison results. The resource comparison results include a judgment result on whether the multimodal data matches, and the matched multimodal data.

[0125] In some exemplary embodiments of this disclosure, the sub-model processing unit may further include a sub-model retrieval unit, which can be used to perform cross-modal retrieval of multimodal data based on the multimodal sub-model used for resource retrieval, and obtain resource retrieval results, including cross-modal retrieval results of multimodal data.

[0126] In some exemplary embodiments of this disclosure, the sub-model comparison unit may include an associated data comparison unit, which can be used to perform data cross-validation on multiple pairs of multimodal data that are related to each other based on the multimodal sub-model used for resource comparison, so as to obtain the associated resource comparison results.

[0127] In some exemplary embodiments of this disclosure, the sub-model comparison unit may further include a historical data comparison unit, a no-difference result feedback unit, and a comparison result feedback unit. Wherein:

[0128] The historical data comparison unit can be used to compare current multimodal data with historical multimodal data;

[0129] The no-difference result feedback unit can be used to directly feed back the no-difference resource comparison result if the current multimodal data is no different from the historical multimodal data.

[0130] The comparison result feedback unit can be used to perform cross-validation of the multimodal data based on the multimodal sub-model used for resource comparison if there are differences between the current multimodal data and the historical multimodal data, so as to obtain the resource comparison result.

[0131] In some exemplary embodiments of this disclosure, the sub-model retrieval unit may include an associated data retrieval unit, which can be used to perform cross-modal retrieval on multiple pairs of interrelated multimodal data based on the multimodal sub-model used for resource retrieval, and obtain associated resource retrieval results.

[0132] In some exemplary embodiments of this disclosure, the sub-model processing unit may further include a sub-model output unit and a decision model output unit. Wherein:

[0133] The sub-model output unit can be used to input multimodal data into a target type multimodal sub-model to obtain the sub-model output result of the target type multimodal sub-model;

[0134] The decision model output unit can be used to input the output results of the sub-model into the decision model, and obtain the processing results of multimodal data according to the decision strategy of the decision model.

[0135] The specific details of each module / unit in the above-mentioned field maintenance device have been described in detail in the corresponding method embodiment section, and will not be repeated here.

[0136] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to exemplary embodiments of this disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0137] Figure 6 A schematic diagram of the structure of a computer system suitable for implementing the embodiments of the present disclosure is shown.

[0138] It should be noted that, Figure 6 The computer system 600 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.

[0139] like Figure 6As shown, the computer system 600 includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 602 or programs loaded from storage section 608 into random access memory (RAM) 603. The RAM 603 also stores various programs and data required for system operation. The CPU 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0140] The following components are connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. A removable medium 611, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on drive 610 as needed so that computer programs read from it can be installed into storage section 608 as needed.

[0141] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 609, and / or installed from removable medium 611. When the computer program is executed by central processing unit (CPU) 601, it performs various functions defined in the system of this disclosure.

[0142] Exemplary embodiments of this disclosure also provide a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the above-described field maintenance method.

[0143] In one embodiment, the computer program product can be a tangible product containing a computer program, such as a computer-readable storage medium storing the computer program. The readable storage medium can be a storage medium based on electrical, magnetic, optical, electromagnetic, infrared, or other signals, including but not limited to: random access memory (RAM), read-only memory (ROM), magnetic tape, floppy disk, flash memory, hard disk drive (HDD), solid-state drive (SSD), etc. For example, the computer program product can be implemented as a non-volatile storage medium storing the computer program, such as read-only memory, NAND flash memory, etc.

[0144] In one implementation, the computer program product can be an intangible product containing a computer program. For example, the computer program product can be implemented as a virtual digital product, such as an executable file, installation package, or other digital file storing the computer program.

[0145] Computer program code can be written in one or more programming languages. Examples of programming languages ​​include C, Java, and C++. Program code can execute entirely on the user's computing device, partially on the user's computing device, or as a standalone software package. It can also execute partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, such as a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via an internet connection provided by a mobile network operator).

[0146] Computer programs can be carried or transmitted via signals such as electrical, magnetic, optical, electromagnetic, and infrared rays. Electronic devices can convert signals carrying computer programs into digital signals, thereby running the computer programs. When a computer program runs on an electronic device, its code is used to cause the electronic device to execute (more specifically, the processor of the electronic device to execute) the method steps of various exemplary embodiments of this disclosure, such as the field maintenance method described above.

[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0148] It should be noted that although several modules for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to embodiments of this disclosure, the features and functions of two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided and embodied by multiple modules.

[0149] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein.

[0150] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A field maintenance method, characterized in that, include: Acquire multimodal data during on-site maintenance; According to the multimodal data processing method, resource identification, resource comparison, or resource retrieval are performed on the multimodal data to obtain the processing result of the multimodal data; On-site maintenance is performed based on the processing results of the multimodal data.

2. The on-site maintenance method according to claim 1, characterized in that, The multimodal data includes at least two of the following: image data, video data, text data, and structured data.

3. The on-site maintenance method according to claim 1, characterized in that, The step of performing resource identification, resource comparison, or resource retrieval on the multimodal data according to the multimodal data processing method to obtain the processing result of the multimodal data includes: The target task is input into the multimodal scene selection model. Based on the output of the multimodal scene selection model, a multimodal sub-model of the target type corresponding to the target task is determined from multiple types of multimodal sub-models. The types of multimodal sub-models include multimodal sub-models for resource identification, multimodal sub-models for resource comparison, and multimodal sub-models for resource retrieval. The multimodal data is processed according to the multimodal sub-model of the target type to obtain the processing result of the multimodal data, which includes resource identification result, resource comparison result, or resource retrieval result.

4. The on-site maintenance method according to claim 3, characterized in that, When the target task is a resource identification task, the step of processing the multimodal data according to the multimodal sub-model of the target type to obtain the processing result of the multimodal data includes: The multimodal data is aligned and / or fused according to the multimodal sub-model for resource identification to obtain the resource identification result, which includes the classification and identification results of the resource type, resource occupancy and resource status of the multimodal data.

5. The on-site maintenance method according to claim 3, characterized in that, When the target task is a resource comparison task, the step of processing the multimodal data according to the multimodal sub-model of the target type to obtain the processing result of the multimodal data includes: The multimodal data is cross-validated according to the multimodal sub-model used for resource comparison to obtain the resource comparison result. The resource comparison result includes the judgment result of whether the multimodal data matches, and the matched multimodal data.

6. The on-site maintenance method according to claim 3, characterized in that, When the target task is a resource retrieval task, the step of processing the multimodal data according to the multimodal sub-model of the target type to obtain the processing result of the multimodal data includes: The multimodal data is subjected to cross-modal retrieval based on the multimodal sub-model used for resource retrieval to obtain the resource retrieval results, which include the cross-modal retrieval results of the multimodal data.

7. The on-site maintenance method according to claim 5, characterized in that, The step of performing data cross-validation on the multimodal data based on the multimodal sub-model used for resource comparison to obtain the resource comparison result includes: Based on the multimodal sub-model for resource comparison, cross-validation is performed on multiple pairs of interrelated multimodal data to obtain the associated resource comparison results.

8. The on-site maintenance method according to claim 5, characterized in that, The step of performing data cross-validation on the multimodal data based on the multimodal sub-model used for resource comparison to obtain the resource comparison result includes: Compare the current multimodal data with historical multimodal data; If the current multimodal data is identical to the historical multimodal data, then the resource comparison result indicating no difference is directly returned. If the current multimodal data differs from the historical multimodal data, then the multimodal data is cross-validated according to the multimodal sub-model used for resource comparison to obtain the resource comparison result.

9. The on-site maintenance method according to claim 6, characterized in that, The step of performing cross-modal retrieval on the multimodal data based on the multimodal sub-model used for resource retrieval to obtain the resource retrieval results includes: Based on the multimodal sub-model for resource retrieval, cross-modal retrieval is performed on multiple pairs of interrelated multimodal data to obtain the associated resource retrieval results.

10. The on-site maintenance method according to claim 3, characterized in that, The step of processing the multimodal data according to the multimodal sub-model of the target type to obtain the processing result of the multimodal data includes: The multimodal data is input into the multimodal sub-model of the target type to obtain the sub-model output result of the multimodal sub-model of the target type; The output of the sub-model is input into the decision model, and the processing result of the multimodal data is obtained according to the decision strategy of the decision model.

11. A field maintenance device, characterized in that, include: The data acquisition module is used to acquire multimodal data during on-site maintenance. The data processing module is used to perform resource identification, resource comparison, or resource retrieval on the multimodal data according to the multimodal data processing method, and obtain the processing result of the multimodal data; The field maintenance module is used to perform field maintenance based on the processing results of the multimodal data.

12. An electronic device, characterized in that, include: processor; as well as A memory for storing one or more programs that, when executed by the processor, cause the processor to implement the field maintenance method as described in any one of claims 1 to 10.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the field maintenance method as described in any one of claims 1 to 10.