A communication scheduling large-screen monitoring method and system based on large model intelligent agent

Through the large model agent predicting the user's question link and updating the communication scheduling screen display interface, the problem of large amounts and types of communication optical cables is solved, and efficient and convenient optical cable maintenance is achieved.

CN119519831BActive Publication Date: 2025-08-12CHINA SOUTHERN POWER GRID COMPANY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411622067.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-08-12
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

The data volume and types of communication optical cables are large, making it difficult for users to quickly and conveniently retrieve the required data from the communication scheduling screen, resulting in low maintenance efficiency of optical cables.

Method used

The communication scheduling large-screen monitoring method based on large-model agents is adopted, and the user's question link is predicted using the link prediction model, and the layout data is generated through the visual language model to update the display interface of the communication scheduling screen in real time, including monitoring data, environmental data, fault data and emergency data.

Benefits of technology

Before the user asks questions, the relevant detailed data will be presented in a timely manner to save interaction time, improve the efficiency and convenience of optical cable maintenance, and avoid the problem from worsening or amplification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119519831B_ABST
    Figure CN119519831B_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a communication scheduling large-screen monitoring method and system based on a large-model intelligent agent, the method comprising: using a historical question data set to conduct supervised training on a link prediction model; obtaining the current question content of the user on the communication optical cable, and using the trained link prediction model to obtain the predicted question link of the current question content; generating prompt content based on the predicted question link, current display data and prompt template, inputting the prompt content into the trained visual language model, and obtaining layout data generated by the visual language model; wherein the layout data includes data to be displayed and display parameters; the data to be displayed is related to the predicted question link, including monitoring data, environmental data, fault data, and emergency data; and updating the display interface of the communication scheduling screen based on the layout data. Provide solutions before the user asks a question to avoid the problem from worsening or expanding. Using a visual language model to manage the layout of the communication scheduling screen can improve management efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of screen display technology, and more specifically, to a communication scheduling large-screen monitoring method and system based on a large model intelligent agent. Background Art

[0002] With the rapid development of communication technology, optical communication networks are one of the most important communication networks in the current power industry. As the scale of optical communication networks continues to expand, maintaining communication optical cables deployed in different geographical areas is a challenging and important task.

[0003] In the prior art, maintenance personnel for optical communication cables can view various data, including the operating status and fault information, on a communication dispatch screen. However, the amount of data involved in optical communication cables is enormous and diverse. Users unfamiliar with the communication system struggle to quickly retrieve the required data for display, and it's also difficult to quickly retrieve the required data from the vast array of data displayed on the dispatch screen. This makes it difficult for users to quickly and conveniently perform optical cable maintenance using the dispatch screen. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a communication scheduling large-screen monitoring method and system based on a large model intelligent agent to solve at least one of the technical problems mentioned above.

[0005] In a first aspect, an embodiment of the present application provides a communication scheduling large-screen monitoring method and system based on a large-model intelligent agent, wherein the large-model intelligent agent includes a visual language model, and the communication scheduling screen is used to display monitoring data, environmental data, fault data, and emergency data of the communication optical cable. The method includes:

[0006] A link prediction model is supervisedly trained using a historical question dataset; wherein the historical question dataset includes multiple historical question links carrying features, each of which includes a historical question content and subsequent question content of a user regarding the display content of a communication scheduling screen; the features include text features and context features of the historical question links;

[0007] Obtaining a current question content of a user regarding the communication optical cable, and using a trained link prediction model to obtain a predicted question link that predicts the current question content, wherein the predicted question link at least includes a next question content regarding the communication optical cable;

[0008] Prompt content is generated based on the predicted question link, current display data of the communication scheduling screen, and a preset prompt template, and the prompt content is input into a trained visual language model to obtain layout data generated by the visual language model; wherein the layout data includes data to be displayed and display parameters; the data to be displayed is related to the predicted question link and includes one or more of the monitoring data, the environmental data, the fault data, and the emergency data;

[0009] A display interface of the communication scheduling screen is updated based on the layout data.

[0010] A second aspect provides an electronic device, including:

[0011] processor;

[0012] a memory for storing processor-executable instructions;

[0013] Wherein, when the processor calls the executable instruction, the method described in the first aspect is implemented.

[0014] A third aspect provides a communication scheduling large-screen monitoring system based on a large model agent, wherein the large model agent includes a visual language model, and the communication scheduling screen is used to display monitoring data, environmental data, fault data, and emergency data of the communication optical cable. The system includes:

[0015] A training module for performing supervised training on a link prediction model using a historical question dataset; wherein the historical question dataset includes multiple historical question links carrying features, each of which includes a historical question content and subsequent question content of a user regarding the display content of a communication scheduling screen; the features include text features and context features of the historical question links;

[0016] a prediction module, configured to obtain the current question content of the user regarding the communication optical cable, and to obtain a predicted question link of the current question content using a trained link prediction model, wherein the predicted question link includes at least the next question content regarding the communication optical cable;

[0017] a layout module, configured to generate prompt content based on the predicted question link, current display data of the communication scheduling screen, and a preset prompt template, and input the prompt content into a trained visual language model to obtain layout data generated by the visual language model; wherein the layout data includes data to be displayed and display parameters; the data to be displayed is related to the predicted question link and includes one or more of the monitoring data, the environmental data, the fault data, and the emergency data;

[0018] An updating module is configured to update a display interface of the communication scheduling screen based on the layout data.

[0019] A fourth aspect provides a computer-readable storage medium having computer instructions stored thereon, which implement the steps of the method described in the first aspect when executed by a processor.

[0020] This application understands and predicts the user's needs by predicting the user's question link, and then uses the visual language model to adjust the layout of the communication scheduling screen based on the user's question link, so that before the user asks the next question, the detailed data related to the user's question can be presented in time according to the predicted next question content, saving the user's time to interact with the system, making the entire process smoother and more efficient, and the display content of the communication scheduling screen can better meet user needs. Providing solutions before the user asks a question helps to avoid the deterioration or expansion of the problem. And using the visual language model to manage the layout of the communication scheduling screen can make the management of the communication scheduling screen more efficient and convenient. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0022] Figure 1 A flow chart of a communication scheduling and large-screen monitoring method based on a large model agent provided in an embodiment of the present application;

[0023] Figure 2 A flow chart of another communication scheduling and large-screen monitoring method based on a large model agent provided in an embodiment of the present application;

[0024] Figure 3 A hardware structure diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0026] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.

[0027] In related technologies, the sheer volume and variety of data involved in optical communication cables makes it difficult for users to quickly and conveniently perform cable maintenance using the communication scheduling screen. To address this issue, the present application provides a large-scale monitoring method for communication scheduling based on a large-scale intelligent agent. The large-scale intelligent agent includes a visual language model, and the communication scheduling screen is used to display one or more of the following: monitoring data, environmental data, fault data, and emergency data for optical communication cables.

[0028] As an example, the communication scheduling screen displays the monitoring data of the communication optical cable, and the display content specifically includes but is not limited to the length of the optical cable, fiber core utilization rate, maintenance level, asset information, and business carrying status. The display screen may include but is not limited to the real-time status and historical trend charts of the communication optical cable, as well as detailed information of important optical cables. The charts include, but are not limited to, line charts, bar charts, etc. The detailed information includes, but is not limited to, fiber core utilization rate and maintenance level, and the detailed information of important optical cables can be highlighted. The monitoring data can also be displayed according to preset display rules, which include giving priority to displaying high-load optical cables according to fiber core utilization rate and / or maintenance level, and using color coding to identify optical cables with different load status. For example, green indicates normal load and red indicates high load.

[0029] As an example, the communication scheduling screen displays the environmental data of the communication optical cable, and the environmental data includes the dynamic environmental data of the environment in which the communication optical cable is located, and the health data of the equipment related to the communication optical cable. Among them, the dynamic environmental data specifically includes but is not limited to ambient temperature, ambient humidity and voltage. The display screen may include but is not limited to real-time environmental parameter charts. The charts include, for example, but are not limited to temperature curve charts, humidity bar charts, voltage line charts, etc. Optionally, abnormal environmental parameters can be highlighted, and detailed abnormal causes and handling suggestions can be provided. Dynamic environmental data can also be displayed according to preset display rules. The display rules include, for example, triggering an alarm based on a preset environmental data threshold, and giving priority to displaying dynamic environmental data that exceeds the environmental data threshold. The display rules also include, for example, using a heat map to display the temperature distribution, and combining the map to display the location and status of the device.

[0030] In addition, health data specifically includes, but is not limited to, device health, the number of normal devices, and the number of abnormal devices. Displays may include, but are not limited to, detailed information and treatment recommendations for abnormal devices, and a device health status chart. The device health status chart displays the ratio of normal to abnormal devices. Health data can also be displayed according to pre-set display rules. For example, these rules may include prioritizing devices with declining health and using color coding to identify abnormal devices.

[0031] As an example, the communication dispatch screen displays the emergency data of the communication optical cable, and the emergency data includes emergency event monitoring data and emergency plan data. Among them, the emergency event monitoring data specifically includes but is not limited to security records, real-time statistical data, maintenance data and emergency dispatch personnel information. The display screen may include but is not limited to the real-time status and processing progress chart of the emergency event, as well as the location and task status of the emergency dispatch personnel displayed using maps and icons. The emergency event monitoring data can also be displayed according to preset display rules. The display rules include giving priority to display according to the severity and scope of the emergency event, and using colors and / or icons to identify emergency events and personnel locations in different states.

[0032] In addition, emergency plan data specifically includes but is not limited to the content of the emergency plan and the status of its execution. The display screen may include but is not limited to detailed information of the currently effective emergency plan and a chart of the plan's execution progress. Optionally, the detailed information of the currently effective emergency plan can be displayed in the form of a card or list. The plan execution progress chart can update the execution status in real time. Emergency plan data can also be displayed according to preset display rules, which include prioritizing the display of key plans and execution status, and using color coding to identify the plan status.

[0033] Of course, the communication scheduling screen is not limited to displaying the data listed above. In actual application, the user can also display other data of the communication optical cable according to actual needs. This application does not limit the data type displayed on the communication scheduling screen.

[0034] Based on this, the monitoring method provided by this application may include the following: Figure 1 Steps 110 to 130 are shown.

[0035] Step 110: Obtain the current question content of the user regarding the communication optical cable, and use the trained link prediction model to obtain a predicted question link for the current question content, wherein the predicted question link at least includes the next question content regarding the communication optical cable.

[0036] Regarding the user's current question about the communication optical cable, as an example, the user can input the current question through a human-computer interaction component. The human-computer interaction component includes, but is not limited to, a keyboard, a mouse, a microphone, etc. The current question can be in text or voice form, which is not limited in this application.

[0037] After obtaining the current question, a link prediction model can be used to predict the subsequent questions the user may ask. The link prediction model training process will be discussed below. These subsequent questions may include one or more. This will yield a predicted question link that includes the current question and at least the next question.

[0038] Step 120: Generate prompt content based on the predicted question link, the current display data of the communication scheduling screen and the preset prompt template, and input the prompt content into the trained visual language model to obtain layout data generated by the visual language model.

[0039] The layout data includes data to be displayed and display parameters; the data to be displayed is related to the predictive question link and includes one or more of the monitoring data, the environmental data, the fault data, and the emergency data.

[0040] Exemplarily, the visual language model can be a large model. The large model refers to a deep learning model trained based on massive text data, and the number of its parameters can be as high as millions or even hundreds of millions. Therefore, the large model can process more complex and comprehensive data, and learn more patterns and rules from it. In the large model, the role of the prompt is to prompt the large model with the context of the input information and the parameter information of the input model. The prompt template consists of instructions, context, examples, input and output. In the prompt word project, the prompt can adopt the RTF prompt word framework. The RTF prompt word framework covers three main elements: role, task and format to ensure that the model generates results that meet user needs.

[0041] In some embodiments, the visual language model used in this embodiment is, for example, a ScreenAI model. Thus, the predicted question link predicted in step 110 and the current display data of the communication scheduling screen can be input into the prompt template of the ScreenAI model to obtain prompt content. The current display data can be a screenshot of the current display interface and / or HTML data corresponding to the current display interface. Subsequently, the prompt content carrying the predicted question link is input into the trained ScreenAI model to obtain layout data generated by the ScreenAI model. It can be seen that the layout data is generated based on the predicted question link prediction. The layout data includes at least the data to be displayed. The data to be displayed includes one or more of the monitoring data, environmental data, fault data, and emergency data mentioned above. Furthermore, in addition to the data to be displayed, the layout data may also include display parameters. Display parameters include, for example, but are not limited to, the display position, size, color, brightness, contrast, animation parameters, and display time of the display content in the communication scheduling screen. By adjusting the display parameters, more important display content can be displayed separately from less important display content.

[0042] Exemplarily, the layout data specifically includes hierarchical information of the display interface in the communication scheduling screen. The hierarchical information is used to represent the layout of the display interface, including information such as element size, position, and color. For example, the hierarchical information may include, but is not limited to, DOM (Document Object Model) elements of the display interface.

[0043] Furthermore, the monitoring data, environmental data, fault data, and emergency data included in the data to be displayed are underlying data stored or generated in real time within the communication system. The layout data may indicate the location and display format of one or more of the monitoring data, environmental data, fault data, and emergency data within the display interface. Examples of such formats include, but are not limited to, charts, maps, and text. Based on the instructions in the layout data, the communication system may read the corresponding data from the storage space and generate the corresponding display interface.

[0044] Optionally, the currently displayed data of the communication scheduling screen includes at least the statistical graph currently displayed in the communication scheduling screen. The statistical graph includes but is not limited to a bar chart, a bar chart, a pie chart, and the like. As an example, after obtaining the current question content of the user in step 110, in addition to using the current question content to predict the question link, it is also possible to determine whether the currently displayed data of the communication scheduling screen includes the question data included in the current question content. The question data includes, for example, but is not limited to, one or more of data about statistical graphs, data about tables, and data about text. If not included, the question data is also included in the data to be displayed determined by the visual language model.

[0045] Step 130: Update the display interface of the communication scheduling screen based on the layout data.

[0046] Finally, based on the visual language model, for example, the layout data output by the ScreenAI model updates the display interface of the communication scheduling screen in real time. In this way, for the display interface of the communication scheduling screen, this embodiment provides real-time update and dynamic interaction functions. Specifically, visualization tools (such as D3.js, ECharts, etc.) can be used to realize the dynamic display of data and the interaction function with the user. The user can click to view the detailed data of the content displayed on the screen, and can also dynamically adjust the display elements and order in the display interface according to user questions and needs.

[0047] As can be seen, this embodiment understands and predicts the user's needs by predicting the user's question chain, and then uses the visual language model to adjust the layout of the communication scheduling screen based on the user's question chain. This allows detailed data related to the user's question to be presented in a timely manner based on the predicted next question content before the user asks the next question, saving the user's time in interacting with the system, making the entire process smoother and more efficient, and the display interface of the communication scheduling screen more in line with user needs. Providing solutions before the user asks a question helps prevent the problem from worsening or expanding. Using the visual language model to manage the layout of the communication scheduling screen can make the management of the communication scheduling screen more efficient and convenient.

[0048] The following is a detailed introduction to steps 110 to 130.

[0049] Regarding the prediction process of the question link in step 110, in some embodiments, a trained link prediction model can be used for prediction. The link prediction model includes, but is not limited to, a sequence model and a Transformer model. Sequence models include, for example, LSTM models and GRU models. Transformer models include, for example, BERT (Bidirectional Encoder Representations from Transformers) models and GPT (Generative Pre-Trained Transformer) models.

[0050] The following is the training process of the link prediction model: Specifically, the user's historical questions and corresponding subsequent questions are first constructed into a historical question dataset. A historical question and its subsequent questions form a historical question link. The historical question dataset includes multiple historical question links.

[0051] Optionally, the data in the historical question dataset can be cleaned and preprocessed. This preprocessing includes, but is not limited to, noise removal, format standardization, and word segmentation. Data cleaning and preprocessing can be performed on a per-link basis. Alternatively, data cleaning and preprocessing can be performed on a per-link basis based on the question content within a historical question link.

[0052] Subsequently, feature extraction can be performed on the historical question link. The feature extraction includes text feature extraction and context feature extraction. Specifically, text feature extraction includes using natural language processing technology to perform feature extraction on the historical question link, and / or perform feature extraction on the question content in the historical question link. For example, one or more technologies in TF-IDF (Term Frequency–Inverse Document Frequency), word vector model and BERT model are used to generate a vector representation of the question link, and / or a vector representation of the question content in the historical question link. The word vector model includes, for example, the Word2Vec model and the GloVe (Global Vectors for Word Representation) model. In addition, the context features include but are not limited to features such as the order of historical questions and the time interval between questions.

[0053] Subsequently, the link prediction model can be supervisedly trained using the historical question dataset. Specifically, the link prediction model's input is a portion of the historical question chain, such as one or more previous questions in the chain, and the model output is the next possible question. By comparing the model's output of the next question with the actual next question in the chain, a loss function is constructed and used to optimize the model parameters, ultimately resulting in a trained link prediction model. Furthermore, during this training process, the historical question dataset used can include extracted features to help improve the link prediction model's prediction accuracy and training efficiency.

[0054] After obtaining a trained link prediction model, when executing step 110, the user's current question can be input into the trained link prediction model. Alternatively, feature extraction can be performed on the current question, including text feature extraction and context feature extraction. The current question, including text features and context features, is then input into the trained link prediction model. The link prediction model can predict the next question based on the current question, obtaining a predicted question link including the current question and the next question. Furthermore, the prediction results of the link prediction model can be used as part of the prompt for training the visual language model.

[0055] Optionally, the next question content predicted by the model can be input into the link prediction model to obtain the next question content, and so on. In this way, a predicted question link including the current question content and multiple subsequent question contents can be obtained.

[0056] Furthermore, in some embodiments, the output of the link prediction model may include multiple candidate question contents. Each candidate question content carries a probability and / or a relevance score, wherein the probability is used to indicate the likelihood that the candidate question content is the next question content after the current question content, and the relevance score is used to indicate the degree of relevance between the candidate question content and the current question content.

[0057] As an example, the candidate question content with the highest probability and / or the highest relevance score among all candidate question contents may be used as the next question content.

[0058] As an example, multiple target candidate question contents with a probability greater than a preset probability threshold and / or a relevance score greater than a preset score threshold can be determined, a predicted question link corresponding to each target candidate question content can be obtained, and subsequent layout data generation can be performed based on multiple predicted question links.

[0059] In addition, optionally, a plurality of candidate question contents may be displayed to the user through the communication scheduling screen. The user may select the question he or she wants to ask from the displayed plurality of candidate question contents, so that the user can ask questions.

[0060] Regarding the fine-tuning process of the visual language model, in order to enable the visual language model to have the ability to generate layout data corresponding to the question content, in some embodiments, the training data of the visual language model may include historical question content carrying historical layout data, so that the visual language model can learn from the training data how to generate corresponding layout data based on the question content. The historical layout data may include screenshots of the historical display interface of the communication scheduling screen, and / or HTML data corresponding to the historical display interface. In addition, in some embodiments, the fine-tuning process of the visual language model may also include the following: Figure 2 Steps 210 to 230 are shown.

[0061] Step 210: Obtain the user's historical behavior data.

[0062] The historical behavior data includes multiple rounds of historical interaction data between the user and the communication scheduling screen, as well as the user's historical question link data. The historical interaction data may include each operation, click, input question and other interactive behavior data of the user when using the user interface (UI) of the communication scheduling screen, as well as the user's stay time data on different pages and functional modules. Optionally, the collected historical behavior data can be cleaned and preprocessed to remove invalid or noise data, and the historical behavior data can be converted into a format suitable for visual language model processing.

[0063] Step 220: Add a low-rank adaptive LoRA layer to the visual language model.

[0064] For example, adding a LoRA (Low-Rank Adaptation) layer typically involves adding additional weight matrices to the linear layer and / or attention layer of the visual language model.

[0065] Step 230: Use the historical behavior data to train the LoRA layer, optimize the parameters of the LoRA layer based on a preset loss function and optimizer, and obtain a trained visual language model.

[0066] The loss function can be selected according to actual needs. The optimizer is used to train the parameters of the LoRA layer, for example including but not limited to the AdamW optimizer.

[0067] The following uses the ScreenAI model as an example to expand the fine-tuning process of the visual language model. First, a data storage mechanism is established to record multiple rounds of user interactions with the communication scheduling screen and data on the question link. Subsequently, the user's historical behavior data is collected based on the data storage mechanism, and the accumulated historical behavior data is used to fine-tune the ScreenAI model. Specifically, a suitable fine-tuning strategy can be determined based on the architecture of the ScreenAI model and the available tools, such as adjusting the parameters of certain layers of the model, adding new layers, or using specific optimization algorithms. In this embodiment, it is chosen to add a LoRA layer to the ScreenAI model, which usually involves adding additional weight matrices to the linear layer and / or attention layer of the model. Define the loss function and optimizer, specifically, determine the evaluation indicators for predicting the user's next question, such as accuracy, recall rate, or F1 value (F1 Score), etc. And construct an objective function based on the evaluation indicators to optimize the model during the training process. The optimizer is used to train the parameters of the LoRA layer, for example, including but not limited to the AdamW optimizer.

[0068] Subsequently, the LoRA layer can be trained using historical behavioral data to adjust its parameters. For example, standard backpropagation and gradient descent can be used to train the model, allowing it to learn the underlying patterns and regularities between the questions. Experiment with different hyperparameters, such as the learning rate, number of training rounds, and regularization parameters, to find the optimal model configuration. Performance metrics, such as loss and accuracy on the validation set, are recorded during training. The performance of the fine-tuned ScreenAI model can then be evaluated on the test set. If performance is poor, the evaluation results can be used to analyze the underlying issues, such as insufficient data, an inappropriate model architecture, or poor hyperparameter settings. Appropriate improvements can be made, such as adjusting hyperparameters, adjusting the model structure, or collecting more user behavior data and re-fine-tuning, until satisfactory results are achieved.

[0069] After fine-tuning the model, you can save the LoRA layer parameters and integrate the fine-tuned ScreenAI model into the communication system, configuring the initial UI layout and related settings. In this way, by continuously optimizing the model parameters, the ScreenAI model can learn user behavior patterns and preferences to improve prediction accuracy.

[0070] In addition, in some embodiments, the prompt template further includes associations between multiple data; wherein the associations are pre-obtained by mining multiple rounds of interaction data between the user and the communication scheduling screen using association rules; the associations include an association between the monitoring data of the communication optical cable and the environmental data, an association between the fault data and the emergency data, and an association between the monitoring data of the communication optical cable and the fault data;

[0071] The generated prompt content carries the association relationship, so that the layout data generated by the visual language model includes first data to be displayed related to the predicted question link, and second data to be displayed related to the first data to be displayed based on the association relationship; wherein the association between the first data to be displayed and the second data to be displayed is represented in the form of a chart and / or a map.

[0072] Association rules are an important technology in data mining, used to discover associations between various data. Association rules may include, but are not limited to, the Apriori (association rule) algorithm and the FP-Growth algorithm.

[0073] When establishing associations between the aforementioned data types, we can first collect user interaction data, i.e., data on multiple rounds of interaction between users and the communication scheduling screen. Optionally, we can cleanse this user interaction data to remove invalid and erroneous data. Furthermore, we can transform this user interaction data to make it suitable for association rule mining.

[0074] Subsequently, association rules, such as the Apriori algorithm, can be used to mine the processed user interaction data. Association rules are generated by setting minimum support and minimum confidence thresholds. These generated association rules can then be evaluated to identify those that are relevant to the communication cable scheduling scenario. Optionally, the mined association rules can be displayed in a chart, text, or other format for easier understanding and application.

[0075] For example, association rule mining can reveal relationships between the aforementioned data types, including but not limited to relationships between monitoring data and environmental data for optical fiber cables, relationships between fault data and emergency response data, and relationships between monitoring data and fault data for optical fiber cables. For example, when a user views a fault on an optical fiber cable, they will likely also view related information charts, including the cause of the fault, fault location, common alarm types, emergency response plans, and safeguards. Therefore, the mined relationships provide valuable guidance for displaying detailed content based on user interaction.

[0076] Since the prompt template includes the association relationship between multiple data, the generated prompt content also carries the association relationship. Inputting the prompt content into the visual language model can enable the visual language model to predict the first data to be displayed based on the predicted predicted question link, and predict the second data to be displayed that has an association relationship with the first data to be displayed. The association between the first data to be displayed and the second data to be displayed is represented in the form of a chart and / or a map. The chart is, for example, a line chart, a bar chart, a pie chart, etc. The map is, for example, a heat map, a map with superimposed status data, a map with superimposed fault data, a map with superimposed emergency data, etc., so that the content displayed on the final communication scheduling screen can better meet user needs.

[0077] The communication scheduling screen displays the correlation between the monitoring data of the communication optical cable and the environmental data. Specifically, the communication scheduling screen displays the following data: ① Temperature and optical cable status line graph, which is used to show the relationship between the ambient temperature change and the optical cable core utilization rate and maintenance level; ② Humidity and optical cable status scatter plot, which is used to show the changes in the optical cable status under different ambient humidity levels; ③ Regional temperature heat map, which shows the temperature distribution of different areas on the map; ④ Optical cable status overlay map, which superimposes the optical cable status identification on the heat map, for example, using color coding to indicate the fiber core utilization rate and maintenance level.

[0078] Fault information and emergency data are closely linked. When an emergency occurs, timely processing of fault information and dispatch of emergency personnel are crucial. The communications dispatch screen displays the fault data of the communication optical cable in a linked manner with emergency data. Specifically, the communications dispatch screen displays the following data: ① A fault handling progress bar, which displays the progress of the fault handling process after the fault occurs, including the dispatch of emergency personnel; ② An emergency response time bar chart, which displays the average response time for different fault types; and ③ A fault location and emergency personnel location map, which displays the fault location and the location and movement of emergency personnel on a map.

[0079] The communication scheduling screen displays the monitoring data and fault data of the communication optical cable in a correlated manner. Specifically, the communication scheduling screen displays the following data: ① A line graph of fiber core utilization and fault probability, which is used to show the relationship between fiber core utilization and fault probability; ② A graph of optical cable status and fault statistics, which is used to show fault statistics under different optical cable statuses; ③ A high-load warning, including when the fiber core utilization exceeds the preset utilization threshold, such as 80%, the high-load area on the map turns red, and the status of the relevant optical cable is marked. At the same time, the fault information of the optical cable is updated in real time, and the fault handling progress and warning information are displayed.

[0080] This embodiment can realize the association display between data, laying the foundation for subsequent warning content stratification and dynamic update.

[0081] In addition, based on any of the above embodiments, the prompt content further includes one or more prompt information, so that the layout data generated by the visual language model also includes third display data obtained based on the prompt information;

[0082] As an example, the prompt information includes prompt information for displaying data related to the current question content mined from a preset database. In this way, in addition to considering the predicted question link, the visual language model also uses data mining techniques to extract data related to the current question content from the preset database as third content to be displayed on the communication scheduling screen.

[0083] As an example, the prompt information includes prompt information for prompting to call a multimodal pre-training model to obtain fused data. Specifically, the text, image, video and other data displayed historically on the communication scheduling screen can be collected, and the data can be pre-processed and labeled so that different types of data can be effectively associated. Subsequently, a multimodal pre-training model, such as a CLIP (Contrastive Language-Image Pre-training) model is used for training to map text, image, video and other data to the same vector space, so that the multimodal pre-training model can perform multimodal learning. After obtaining the multimodal pre-training model, in addition to predicting the question link, the visual language model also calls the multimodal pre-training model to fuse the multimodal data related to the current question content obtained from multiple data sources, for example, mapping them to the same vector space, and the obtained fused data is the third display data. In this way, different types of multimodal information related to the current question content can be integrated and displayed on the communication scheduling screen, realizing the integrated display of different data types (including text, images, videos, and sensor data) on the large screen, ensuring that users can see the most relevant content when interacting, and improving the usability of the communication scheduling screen.

[0084] As an example, the prompt information includes prompt information for prompting the call of the knowledge graph. Specifically, the relevant entities of the optical cable and the relationships between the entities can be collected, and the knowledge graph of the optical cable can be constructed using multiple entities and their relationships. The entities include, for example, but are not limited to, optical cable types, fault types, maintenance steps, etc. In this way, in addition to predicting the question link, the visual language model also uses an inference engine (such as an OWL inference engine) to process the current question content, and derives data related to the current question content through the relationship chain in the knowledge graph as the third content to be displayed in the communication scheduling screen. In this way, a question-and-answer system can be created that can understand user questions and provide accurate answers, ensuring that users can quickly obtain the required information. It also supports real-time updates of the knowledge graph to ensure that the displayed information is always the latest.

[0085] As another example, the prompt information includes prompt information for prompting the call of the adaptive model. Specifically, professional question and answer data in the field of optical cables can be collected and annotated to ensure that the system can adapt to the needs of different scenarios. The scene adaptive model is then trained using a contrastive learning method so that it can better distinguish between subtle semantic differences in different scenarios. Subsequently, the trained scene adaptive model can be used to automatically adjust the query parsing logic according to the context of the current question content and the question scene, so that the displayed content is more relevant. In this way, the visual language model can obtain the third content to be displayed based on the context of the current question content and / or the question scene by calling the scene adaptive model, so that the displayed content is more relevant. In addition, the system can also use domain adaptation technology and contrastive learning to continuously adjust and optimize the scene adaptive model according to the user's interactive behavior (such as clicks, browsing time, etc.), so that the communication scheduling screen can understand and adapt to changes in user queries in specific optical cable scenarios.

[0086] Ultimately, the aforementioned data mining methods, multimodal pre-trained models, knowledge graphs, and adaptive models can be integrated into a complete communication system as software modules, allowing the visual language model to generate layout data by invoking corresponding auxiliary methods through an API (Application Programming Interface).

[0087] This embodiment is described by the following examples:

[0088] If the user's current question is "What are the current high-priority alarms?", the visual language model will retrieve and display high-priority alarm information. For example, if the prompt includes information that indicates data mined from a pre-set database related to the current question, the visual language model will use data mining methods to quickly extract high-priority alarm information related to the user's query from the vast amount of alarm data. The mining algorithm also evaluates the priority, history, and potential impact of all alarms, and displays this data on the communication scheduling screen.

[0089] For example, if the prompt includes a prompt to call a multimodal pre-trained model to obtain fused data, the visual language model will call the multimodal pre-trained model to fuse multiple modalities of data, such as alarm data, geographic information, and video surveillance, to form a unified display. For example, the location of a high-priority alarm can be displayed using a combination of maps, real-time video, and sensor data.

[0090] As an example, if the prompt content includes prompt information for calling the knowledge graph, the visual language model calls the knowledge graph to understand the relationship between the alarm data and other relevant information (such as historical data, fault causes, treatment solutions, etc.), and provides accurate display data based on these relationships.

[0091] For example, if the prompt includes information that prompts the invocation of the adaptive model, the visual language model will automatically adjust the query parsing logic based on the user's alert, making the results more relevant. For example, if a user queries "Where is the latest alert?", the system will prioritize displaying the geographic location.

[0092] By combining various prompts, we can present relevant and detailed data based on the user's current question, allowing users to quickly and easily access comprehensive information related to their question. This has built a highly intelligent visual language model system that accurately understands and presents user needs, making communication scheduling screen management more efficient and convenient.

[0093] Furthermore, the communication scheduling screen provided in this application also features scene adaptation. Specifically, the visual language model can adjust the display data of the communication scheduling screen to suit different scenarios. Thus, based on any of the above embodiments, the prompt content also includes the current scene information input by the user; and the data to be displayed is related to the current scene information.

[0094] As an example, the current scenario includes a daily communication scheduling scenario. If the prompt content includes information about the daily communication scheduling scenario, the visual language model will adjust the data to be displayed based on the daily communication scheduling scenario. Specifically, in the daily communication scheduling scenario, the data to be displayed includes monitoring data, environmental data, and fault data.

[0095] For example, in daily communication scheduling scenarios, the status of communication optical cables directly affects communication quality and scheduling efficiency. Displaying the current status of each cable and its historical trends helps identify potential maintenance needs and optimize communication scheduling. Based on this, the monitoring data displayed in daily communication scheduling scenarios includes visualization of communication routes, including cable length, asset level, maintenance level, fiber core utilization, and traffic carrying capacity. Optionally, based on established associations, communication cables with high fiber core utilization and low maintenance levels, which may be prone to failure, can be monitored specifically.

[0096] For example, in daily communication scheduling scenarios, dynamic environmental conditions can affect the performance of communication optical cables. By monitoring environmental data, environmental anomalies can be promptly detected and addressed. Based on this, the environmental data displayed in daily communication scheduling scenarios includes ambient temperature, ambient voltage, ambient humidity, and their changing trends. Optionally, based on established associations, it can be determined that abnormal environmental parameters (such as high temperature and high humidity) can cause changes in optical cable status. Therefore, this data can be displayed in conjunction with monitoring data for communication optical cables.

[0097] For example, in daily communication scheduling scenarios, fault data is closely related to communication cables and environmental conditions. By monitoring communication cable faults, problems can be quickly located and resolved. Based on this, the fault data displayed in daily communication scheduling scenarios includes the fault title, location, cause, type, response time, and processing progress. Optionally, based on established associations, monitoring data and environmental data associated with the fault data of the faulty cable can also be displayed in a linked manner.

[0098] As another example, the current scene includes an emergency command scene. If the prompt includes information about the emergency command scene, the visual language model will adjust the displayed data based on the emergency command scene. Specifically, in the emergency command scene, the displayed data includes optical cable topology data, alarm data within the emergency area, and support data for the emergency area.

[0099] For example, in an emergency command scenario, the optical cable topology data for the emergency area can be displayed as a cable topology map. Alternatively, based on the established associations, the status of the optical cables in the emergency area is closely related to the emergency response, so the monitoring data of the communication cables can be combined with the emergency support data for display.

[0100] For example, in an emergency command scenario, alarm information can help quickly identify the source of the problem. For this purpose, the alarm data displayed in the emergency area can include alarm level, location, time, and more. Optionally, based on the established associations, the alarm information is associated with communication cable monitoring data and emergency support data. Therefore, during emergency response, monitoring data for communication cables in the alarm area and the location of emergency personnel can be displayed together.

[0101] For example, in an emergency command scenario, emergency support records can ensure transparency and traceability of the emergency response, facilitating post-event summary and optimization of emergency plans. For this purpose, support data displayed in the emergency area during the emergency command scenario includes emergency personnel location, dispatch status, and processing progress. Optionally, based on established associations, emergency support data can be displayed in conjunction with fault data and communication cable monitoring data to ensure transparency and traceability of the emergency response.

[0102] As another example, the current scene includes a holiday-themed scene. If the prompt content includes information about the holiday-themed scene, the visual language model will adjust the data to be displayed based on the holiday-themed scene. Specifically, in the holiday-themed scene, the data to be displayed includes: monitoring data of target communication optical cables in the holiday support area, alarm information in the holiday support area, and management information for the holiday support area.

[0103] For example, in a holiday scenario, since holidays are peak communication periods, monitoring the operating status of communication optical cables can ensure stable operation of the communication system under high load. Based on this, in this holiday scenario, the monitoring data displayed for the target communication optical cables in the holiday support area includes the target communication optical cables' cable length, fiber core utilization, and traffic carrying capacity. Optionally, based on the established associations, since the high load during holidays requires special monitoring, the impact of the environment on the status of the communication optical cables can be displayed in conjunction with the displayed environmental data.

[0104] For example, in a holiday-themed scenario, timely processing of alarm information can ensure smooth communication during holidays. Therefore, in this holiday-themed scenario, the alarm information displayed in the holiday guarantee area includes alarm level, location, time, and more. Optionally, based on the established associations, alarm information is associated with monitoring data and fault data of communication optical cables. During holidays, high loads on communication optical cables may lead to frequent alarms. Therefore, this data can be combined with monitoring data, fault data, and alarm data for a focused display.

[0105] For example, in a special holiday scenario, customized templates can improve the efficiency and relevance of holiday support work. For this reason, the management information displayed for the holiday support area in this special holiday scenario can include customized management templates that include common alarm types, emergency plans, support measures, and more. Optionally, based on the established associations, customized templates should be combined with historical data and current monitoring data to ensure efficient and orderly support work during the holidays.

[0106] As another example, the current scenario includes a reporting scenario. The prompt content includes prompt information for the reporting scenario, and the visual language model adjusts the displayed data based on the reporting scenario. Specifically, in the reporting scenario, the displayed data includes: communication optical cable indicator statistics, indicator analysis data, and service assurance status data.

[0107] For example, in a reporting scenario, a summary display provides a comprehensive understanding of the communication system's operational status, facilitating reporting to management and decision-making. For this purpose, the communication optical cable indicator statistics displayed in the reporting scenario include cable status statistics, fault statistics, and emergency response record statistics. Optionally, these indicator statistics can be linked to data to be displayed in daily communication scheduling scenarios, emergency command scenarios, and holiday-themed scenarios to fully demonstrate the system's operational status.

[0108] For example, in a reporting scenario, performance indicator analysis can be used to assess the efficiency and effectiveness of communication system operations, providing data support for subsequent improvements. Therefore, the indicator analysis data displayed in the reporting scenario refers to the statistical analysis results of various performance indicators, including fiber core utilization, fault handling efficiency, and emergency response time. Optionally, performance indicator analysis data can also be displayed in combination with historical data and the current status of the communication optical cable to evaluate the efficiency and effectiveness of system operations.

[0109] For example, in reporting scenarios, service assurance status can directly reflect the stability and reliability of the communication system, enabling timely identification and resolution of issues to ensure normal service operation. Based on this, service assurance data displayed in reporting scenarios includes service volume, service quality, and failure rate. Furthermore, service assurance data can be correlated with communication cable monitoring data, fault data, and environmental data to ensure stable service operation.

[0110] In addition, when faced with the joint display of multiple types of data, the layout of the multiple types of data in the communication scheduling screen can also be determined according to the following rules. The rules include:

[0111] 1) Prioritize display of important information: Important information (such as high-priority alarms, high-load optical cables, and emergency response data) is displayed in the important display area of the communication scheduling screen. The important display area is, for example, the center or top of the screen, and its area on the screen is not less than a preset threshold, so that the user's attention is focused on the important display area.

[0112] 2) Tight display of related data: Various types of data with related relationships are displayed adjacent to each other on the communication scheduling screen, making it easy to compare and analyze.

[0113] 3) Display data in chronological order. For example, data with chronological order such as fault handling progress and emergency response records can be displayed in the form of a timeline or carousel.

[0114] 4) Display data in order of priority. For example, data such as alarm information and fault information can be prioritized and sorted using color coding and lists, with high-priority data displayed first.

[0115] Through the above display logic and display method, input support can be provided for the visual language model, dynamic adjustment and personalized display can be performed, ensuring that the system can flexibly respond to different needs and warning scenarios.

[0116] In addition, in addition to different layouts of the communication scheduling screen according to different scenarios, the visual language model can also perform screen information understanding, question answering, UI navigation and content summary based on the current display content of the communication scheduling screen in different scenarios.

[0117] Taking the ScreenAI model as an example, screen information understanding means that the ScreenAI model can identify and understand the elements of the currently displayed content and the content of infographics on the communication scheduling screen, including their types, locations, and relationships with each other. Question answering means that the ScreenAI model can understand the visual information obtained from the currently displayed content on the communication scheduling screen and answer questions about the displayed content, such as the content of infographics. UI navigation means that the ScreenAI model can interpret the navigation instructions input by the user (such as the "return" instruction) and identify appropriate UI elements from the currently displayed content for interaction, understand user intentions, and be able to navigate accurately in the interface. Content summary means that the ScreenAI model can concisely summarize the currently displayed content, as well as extract and summarize the core points of the currently displayed content.

[0118] In order to enable the visual language model to realize screen information understanding, question answering, UI navigation and content summary functions, the visual language model can be trained using screenshots carrying target information. The target information includes: question and answer pairs about the screenshot, logical relationship information between elements in the screenshot, navigation operation information, and summary information. The specific training process is as follows:

[0119] First, we collected training data. Specifically, we collected screenshots of various UI elements, infographics, and other content displayed on the communication scheduling screen, annotating the location, type, and relationship of each element. We also prepared corresponding question-and-answer pairs for each screenshot, covering the content of the UI elements and infographics in the screenshot.

[0120] Subsequently, computer vision techniques are used to train a visual language model to identify UI elements and infographics on the screen. Specifically, methods such as semantic segmentation and object detection can be used to identify element types and locations. The logical relationship information contained in the screenshots can be used to train the visual language model to understand the logical relationships between elements (for example, the relationship between a button and a text box), enabling the visual language model to understand screen information.

[0121] In addition, natural language processing technology can be used to train visual language models to understand questions, and by combining screen content and question pairs, the visual language model can be trained to generate answers to realize question-answering functions.

[0122] In addition, the navigation operation information carried in the screenshot can train the visual language model to understand natural language instructions, such as the "return" instruction, the "open settings" instruction, etc., so that the visual language model can identify UI elements related to user instructions and realize UI navigation functions.

[0123] Furthermore, natural language generation technology can be used to train a visual language model based on the summary information carried by the screenshot to extract key information from the screen content, so that the visual language model can generate a concise and clear summary, realizing the content summarization function.

[0124] Finally, the trained visual model can be encapsulated as an API interface, and a user interface can be designed to allow users to enter questions or commands. By calling the trained visual language model through the API, screen information understanding, question answering, UI navigation, and content summarization can be performed based on the current display content of the communication scheduling screen in different scenarios.

[0125] Based on this, the user can also enter user instructions, including one or more of screen information comprehension instructions, question and answer instructions, navigation instructions, and summary instructions. In this way, because the prompt content of the visual language model includes the current scene information entered by the user and the user instruction, the visual language model can execute the user instruction on the to-be-displayed data corresponding to the current scene information, including one or more of screen information comprehension, question answering, navigation operations, and summary generation.

[0126] For example, in a daily communications scheduling scenario, if the current display on the communications scheduling screen is monitoring data for communications optical cables, a trained visual language model can be used to understand the screen information. Specifically, the visual language model can identify and understand the visual information of the currently displayed communications optical cables, including cable length, asset level, maintenance level, fiber core utilization, and service load. Furthermore, the visual language model can understand logical relationships, specifically the status of different communications optical cables and their interrelationships, such as the fact that highly utilized cables may require more maintenance.

[0127] Furthermore, trained visual language models can be used to answer questions. For example, if a user asks, "Which optical cable has the highest fiber core utilization rate?", the visual language model can answer based on the currently displayed monitoring data of the communication optical cable, such as, "According to the current data, the fiber core utilization rate of cable A is 85%, the highest among all optical cables."

[0128] Furthermore, a trained visual language model can be used to generate content summaries. For example, based on the currently displayed monitoring data of a communication optical cable, a summary can be generated: "Today, the fiber core utilization rate of two optical cables exceeded 80%, and one of them, Cable B, has a lower maintenance level and requires special attention."

[0129] Furthermore, the trained visual language model can be used for UI navigation. For example, if a user enters the navigation command "Show me the optical cables with the lowest maintenance rating in the last month," the visual language model can parse the command and, based on the currently displayed monitoring data for the communication cables, display a list of optical cables with the lowest maintenance rating in the last month that match the command.

[0130] Taking the emergency command scenario as an example, if the communication dispatch screen currently displays the optical cable topology data for the emergency area, the trained visual language model can be used to understand the screen information. Specifically, the visual language model can automatically identify the nodes, connection lines, and possible fault points in the communication cable topology diagram displayed on the communication dispatch screen. It can also understand the visual elements in the topology diagram and convert them into structured data, such as the location information of the starting point, end point, and intermediate connection points of the optical cable.

[0131] Furthermore, the trained visual language model can be used to answer questions about the status of optical cables based on the optical cable topology data of the currently displayed emergency area. For example, questions such as "Which optical cables are operating normally?" or "Where is the nearest fault point?" can be answered.

[0132] In addition, the trained visual language model can be used to perform content summarization. For example, a concise summary of a communication cable topology map can be generated, highlighting important information such as the status of key nodes and potential risk points.

[0133] Furthermore, the trained visual language model can be used for UI navigation. For example, it can navigate the on-screen topology map based on the user's voice or text navigation instructions, such as zooming in on a specific area or displaying the status of a specific line.

[0134] Taking a holiday-themed scenario as an example, if the communication dispatch screen currently displays alarm information for the holiday support area, the trained visual language model can be used to understand the screen information. Specifically, the visual language model can understand the alarm information displayed on the screen, including the alarm level, location, time, etc., and can interpret color codes and map markers to quickly locate the problem.

[0135] In addition, the trained visual language model can be used to answer questions about alarms based on the alarm information in the currently displayed holiday protection area. For example, questions like "What are the most frequent alarm types during holidays?" or "Where did the latest alarm appear?" can be answered.

[0136] In addition, a trained visual language model can be used to perform content summarization. For example, a concise summary report can be generated, outlining the key points of alarm information during the holiday period, such as the number of alarms, type distribution, and processing progress.

[0137] Furthermore, the trained visual language model can be used for UI navigation. For example, it can navigate to the indicated alarm details according to the user's navigation instructions, such as viewing the handling progress of a specific alarm or displaying the history of alarms at the same location.

[0138] Taking a reporting scenario as an example, if the communication dispatch screen currently displays statistics on communication fiber optic cables, the trained visual language model can be used to understand the information on the screen. Specifically, the visual language model can identify various statistical indicators displayed on the screen, including fiber optic cable status, fault conditions, emergency response records, etc., and understand the logical relationships between these data.

[0139] In addition, the trained visual language model can be used to answer questions about statistics based on the currently displayed communication fiber optic cable indicator statistics. For example, it can answer questions such as "How many fiber optic cable failures have there been this month?" or "What was the average emergency response time last quarter?"

[0140] Furthermore, trained visual language models can be used for content summarization. For example, a concise summary report can be generated that summarizes key statistics for various indicators, such as the stability of optical cable status, fault trends, and the effectiveness of emergency response.

[0141] Furthermore, the trained visual language model can be used for UI navigation. For example, it can navigate to the indicated statistical data view according to the user's navigation instructions, such as switching to view statistics for different time periods or adjusting the chart type to present data more clearly.

[0142] It can be seen that this embodiment uses the visual language model to implement screen information understanding, question answering, UI navigation and content summary functions for different display data in different scenarios, which greatly facilitates users to use and manage the communication scheduling screen.

[0143] Furthermore, in some cases, users' questions may contain colloquial elements, which may cause the visual language model to fail to fully understand the user's intent. To address this issue, this embodiment proposes that the visual language model's prompt content can also include the current question content and question-structuring task information. The visual language model then converts the current question content into a structured question based on the question-structuring task information and generates an answer.

[0144] This embodiment defines a question structuring task, which is used to perform structural analysis on the user's current question content, so that the visual language model can understand the user's intention from the converted structured question and generate answer content.

[0145] Specifically, the question structuring tasks include determining the scope of the question topic, determining the type of data involved in the question, constructing the question framework, refining the question elements, and setting the answer format.

[0146] The so-called determination of the scope of the question topic refers to determining whether the current question content is centered around the communication optical cable related data displayed in the communication scheduling screen.

[0147] Determining the data type involved in the question refers to classifying and organizing the data type corresponding to the current question content. The data type includes optical cable core utilization rate, environmental parameters and their impact on optical cable performance, computer room temperature thresholds and response measures, faulty optical cable information, and unresolved optical cable faults, etc.

[0148] Question framing refers to transforming the current question into one or more of the following: general questions (e.g., "Please explain the specifics of [data category]"), specific numerical questions (e.g., "What is the numerical value of [data category]"), comparative questions (e.g., "How has [data category] changed compared to [time period]"), cause analysis questions (e.g., "What is the cause of the change in [data category]"), impact assessment questions (e.g., "What impact has [data category] had on [relevant aspects]"), prediction questions (e.g., "What will happen to [data category] in the future [time period]"), and solution questions (e.g., "Are measures being taken to address [problem]?"). The content in brackets [] is extracted from the current question.

[0149] The so-called refinement of problem elements refers to the refinement of time, space, objects, conditions, causes and other elements involved in the problem. Specifically, refining time elements means refining the specific time period, such as accurate to the year, month, day, hour, minute, and second, and determining whether it is the time range in the past, present, or future. Refining spatial elements means refining the specific geographical location or regional scope, and determining the specific spatial objects such as places and buildings involved. Refining object elements means clarifying the specific object targeted by the problem, whether it is a person, object, event, or concept, and providing a more detailed description of the characteristics and attributes of the object. Refining condition elements means determining the preconditions, restrictions, or assumptions contained in the problem, such as resource availability, environmental factors, policy regulations, etc.

[0150] Setting the answer format refers to specifying the format of the answer, for example, specifying that the answer content be output in the form of a list, a chart, or a detailed text description.

[0151] For example, here's an example of a prompt that defines a structured task with a problem:

[0152] “Question topic scope: Optical cable related data in the communication monitoring large screen system.

[0153] Data type classification: 1. Fiber optic cable core utilization rate, 2. Environmental parameters and their impact on fiber optic cable performance, 3. Equipment room temperature thresholds and countermeasures, 4. Faulty fiber optic cable information, 5. Unresolved fiber optic cable faults.

[0154] Example: Problem framework: 1. Fiber optic cable core utilization rate

[0155] -General question: Please explain the specific situation of fiber optic cable core utilization rate.

[0156] -Specific numerical question: What is the current value of the optical cable core utilization rate?

[0157] -Comparison question: How has the fiber optic cable core utilization rate changed compared to last week?

[0158] -Cause analysis question: What is the reason for the reduced utilization rate of optical cable cores?

[0159] - Impact assessment question: What impact does excessive fiber core utilization in optical cables have on communication quality?

[0160] -Forecast question: What is the likely rate of fiber optic cable core utilization in the next week?

[0161] -Solutions: Have measures been taken to address the low utilization rate of optical cable cores?

[0162] Refine the problem elements: 1. Time element:

[0163] - Past: For example, "What was the fiber optic cable core utilization rate in August last year?"

[0164] -Now: "What is the temperature in the computer room at this moment?"

[0165] -Future: "Can the faulty fiber optic cable be repaired within the next three days?"

[0166] 2. Spatial elements:

[0167] - Specific geographic location: "What is the humidity in the computer room in area A?"

[0168] - Regional scope: "How many fiber optic cable faults are there in the entire campus?"

[0169] 3. Object elements:

[0170] -Specific: "What is the fiber core utilization rate of a particular fiber optic cable?"

[0171] -Object characteristic attribute: "What is the failure frequency of optical cables with high maintenance level?"

[0172] 4. Conditional elements:

[0173] -Prerequisite: "How to improve the efficiency of optical cable maintenance when resources are limited?"

[0174] - Constraints: "How does the cable perform if the temperature does not exceed 30 degrees Celsius?"

[0175] 5.Cause factors:

[0176] - Direct cause: "What is the direct cause of this fiber optic cable failure?"

[0177] -Root Cause: "What is the root cause of the long-term low utilization rate of optical cable cores?"

[0178] Set format requirements:

[0179] Responses should be presented in a list format, with each question's response including the following:

[0180] 1. Description of the problem, 2. Specific answer, 3. Relevant data or charts

[0181] By converting the current provided content input by the user into a structured question through a question structuring task, it is possible to ensure that the question focuses on the display content of the communication scheduling screen and avoids deviating from the topic. At the same time, the question framework covers different types of inquiries, and the comprehensiveness of the question can be ensured by refining the question elements. And by constructing different types of question frames, the visual language model can better understand the user's question intention. By clarifying elements such as time, space, and objects, it helps the visual language model to give more specific operational suggestions or solutions. By refining the question elements, misunderstandings in communication can be reduced and more effective information exchange can be promoted. In summary, through this embodiment, not only can the accuracy and effectiveness of questions be improved, but also respondents can be encouraged to provide more targeted and practical answers, thereby better serving the management and decision-making process of the communication monitoring large screen system.

[0182] In addition, the visual language model can also layout multiple communication scheduling screens at the same time. And each communication scheduling screen can be located in a different geographical location. For example, the visual language model can layout the communication scheduling screen located in City A and the communication scheduling screen located in City B at the same time. In some possible scenarios, multiple regions may need to collaborate to jointly control communication optical cables. At this time, personnel from each region will jointly control personnel from other regions on their own communication scheduling screens. However, the communication scheduling screens in different regions may display different content, which will cause the messages of personnel in the two regions to be unable to be synchronized, which brings great inconvenience to joint control. The following three scenarios are listed to illustrate:

[0183] Scenario 1: When two communication scheduling screens are collaborating on an alert, one might display a fiber optic cable fault—for example, centered around a diagram depicting the fault. The other might display the current alarm status of the fiber optic cable. In this scenario, one person might focus on the fault details, while the other focuses on the alarm status. If the displays on the two screens are out of sync, this can lead to information mismatches when discussing the fault's impact and alarm handling. Furthermore, if the two individuals discuss based on different displays and fail to proactively reference the other person's screen, critical information may be missed, impacting decision-making accuracy. Furthermore, the discrepancies between the two screens can lead to misjudgments of the urgency of a particular fault when prioritizing action.

[0184] Scenario 2: Two communication scheduling screens facilitate alert processing. Screen A focuses on the transmission network topology and network element alarm information, while Screen B displays pie charts of alarm types and the number of unresolved alarms in different regions. In this scenario, the different regions and alarm types focused on by screen can make it difficult for users to simultaneously understand the overall alert situation. For example, they can't quickly determine the severity and impact of an alarm in a specific region. Furthermore, because Screen A displays the network topology and specific node locations of alarms, while Screen B categorizes alarm information by region, the alarm types and numbers for different regions may appear as statistical pie charts on Screen B, but not as intuitively visible on Screen A, potentially leading to regional information mismatches. Furthermore, Screen A focuses on network element alarm information within the network topology, while Screen B displays alarm severity and the number of unresolved alarms. This separation can lead to user bias in assessing alarm severity, especially when making quick decisions.

[0185] To address the issue of inconsistent displays on multiple communication scheduling screens, this embodiment proposes that a multi-screen display content monitoring task can be input into a prompt template. The prompt content includes the multi-screen display content monitoring task, so that the visual language model performs difference detection on the display content of at least two communication scheduling screens based on the multi-screen display content monitoring task. The difference detection includes time synchronization difference detection, monitoring value difference detection, and image content difference detection. If the difference detection indicates that the display content difference between at least two communication scheduling screens is greater than a preset difference threshold, a synchronization instruction is output to the at least two communication scheduling screens, so that the communication scheduling screens consume, read, and display data from the same message queue.

[0186] To enable the visual language model to detect differences between the display contents of multiple communication scheduling screens and complete data updates, the visual language model can be trained first. The following describes the training process of the visual language model.

[0187] First, a large number of difference data sets of screen display content can be collected, and the difference data sets include time lag type data, numerical difference type data, image difference type data, etc. And each type of difference data includes screen data of two communication scheduling screens and processing operations on the two communication scheduling screens. Among them, the screen data includes timestamps, number of alarms, and image snapshots, and the processing operations include synchronization operations and data update operations, etc. Subsequently, the collected difference data sets can be used to generate a training data set, and the optimal processing method corresponding to each difference type can be marked. Subsequently, the training data set can be used to train the visual language model. In this way, the visual language model can not only understand the differences between the screen display contents during training, but also make optimal processing based on the differences.

[0188] The following are three prompt template examples for training visual language models:

[0189] Example 1: "A significant difference was detected between the images displayed on screens A and B for optical cable t: Screen A displays a bar graph of the service carrying capacity and a graph of the fiber core utilization rate for optical cable t. Screen B displays the global topology and alarm location distribution for optical cable t."

[0190] The synchronization strategy used for processing and the optimal processing method:

[0191] Q1: Do I need to obtain the HTML format of screens A and B and parse the HTML? A1: The visual language model obtains the HTML content of screens A and B through the API, uses JavaScript to parse the HTML content, and extracts the data display area that needs to be synchronized.

[0192] Q2: If the image difference detection SSIM falls below the set threshold of 0.8, is it necessary to output the HTML format of screens A and B and synchronize the image information of screens A and B? A2: Yes. The visual language model calls the API of the backend server to obtain the data that needs to be synchronized. For example, it sends a GET request to the API to obtain the service carrying and alarm data of optical cable t, obtains the data of screen B (such as the alarm distribution map and global topology map) from the API, and injects this data into the corresponding HTML elements of screens A and B.

[0193] Q3: Is it necessary to synchronize the information on screens A and B based on real-time prediction technology? A3: Yes. Based on the operating patterns of operators A and B, it is predicted that operator A may further investigate the alarm location, while operator B may focus on changes in fiber core utilization. Therefore, the information on the two screens is synchronized in advance, and the predicted data is injected into the corresponding HTML elements of screens A and B, allowing screens A and B to display the alarm location and fiber core utilization simultaneously.

[0194] Q4: Is it necessary to use multimodal data fusion technology to simultaneously display fiber core utilization, alarm information, and dynamic environment monitoring data? A4: Yes. The system should fuse fiber core utilization, alarm information, and dynamic environment monitoring data and inject the data into the corresponding HTML elements on both screens for simultaneous display. This allows operators to simultaneously view the fiber cable's high-load operating status, alarm locations, and external environmental impacts on both screens, avoiding misunderstandings.

[0195] Q5: Is it necessary to use the knowledge graph to analyze the correlation between historical alarms and the current status and display the cause of the alarm? A5: Yes. The system should automatically analyze the correlation between the optical cable's historical alarms and the current status, infer the cause of the alarm, and inject the alarm context and related network element status into the corresponding HTML elements on screens A and B to help operators understand the detailed context of the current alarm.

[0196] Through this training process, the trained visual language model is able to handle a variety of screen display content differences, improving its adaptability to different scenarios. This enables the visual language model to perform well in a variety of application scenarios. Furthermore, by designing specific training prompts, the visual language model can more accurately identify and handle differences in display content between screens, reducing the possibility of misunderstandings and incorrect processing, thereby improving the overall accuracy of the system.

[0197] Prompts designed for different types of screens can help visual language models respond more quickly, such as synchronizing data in a timely manner to ensure information consistency, thereby reducing operational problems caused by delays or errors.

[0198] Using the trained visual language model, a multi-screen display content monitoring task can be performed. Exemplarily, this multi-screen display content monitoring task is actually a dynamic error detection task. Based on this multi-screen display content monitoring task, the visual language model can continuously monitor the content displayed on at least two communication scheduling screens and continuously detect differences between their displayed content.

[0199] In order to realize the multi-screen display content monitoring task (dynamic error detection task), first, multi-level error thresholds can be set, such as time lag threshold, numerical difference threshold and image difference threshold, and these detection thresholds can be initialized. Among them, the time lag threshold is defined as the time synchronization error between the two communication scheduling screens, in milliseconds. The numerical difference threshold is used to monitor the difference in numerical information in the display content, such as the number of alarms, the number of fault statistics, etc. The numerical difference threshold can be a percentage error or an absolute numerical error. The image difference threshold can be defined using the structural similarity index (SSIM) or the mean square error (MSE) between the two images.

[0200] After completing the initialization and configuration of the detection threshold, the visual language model can perform the above-mentioned multi-screen display content monitoring task, that is, perform time synchronization difference detection, monitoring value difference detection, and image content difference detection on the display content of at least two communication scheduling screens.

[0201] For example, deep learning real-time monitoring algorithms, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), can be used to continuously monitor the display content of multiple communication scheduling screens. Furthermore, optical character recognition (OCR) technology and image comparison algorithms can be used to regularly capture and analyze display content to detect discrepancies in numerical values, text, or images.

[0202] Specifically, when performing time synchronization difference detection, the update time of each piece of data on multiple communication scheduling screens can be marked with a timestamp. The time difference Δt is periodically calculated based on the timestamp of each communication scheduling screen. The time difference is compared with a preset time lag threshold. If the time difference exceeds the time lag threshold, the data synchronization mechanism is triggered.

[0203] When performing monitoring numerical difference detection, the numerical difference can be characterized by difference calculation or percentage deviation calculation. The difference calculation is, for example, the absolute value of the difference between two values of the same parameter in two communication scheduling screens, denoted as |AB|, where A and B are the values of the same parameter on the two communication scheduling screens. The percentage deviation calculation is, for example, the absolute value of the difference between two values of the same parameter in two communication scheduling screens and the percentage of one of the values, denoted as |(AB) / B|*100%. Subsequently, the numerical difference can be compared with a preset numerical difference threshold. If the numerical difference exceeds the numerical difference threshold, the data synchronization mechanism is triggered.

[0204] When detecting image content differences, snapshots of the display content of two communication scheduling screens can be captured and compared using an image processing algorithm (such as SSIM or MSE). For example, the SSIM output value ranges from [0, 1] and is negatively correlated with the image content difference. That is, the closer the output value is to 1, the smaller the difference in image content between the two snapshots, and the more similar the displayed content. The SSIM value can then be compared with a preset image difference threshold. If the SSIM value falls below the threshold, it is flagged as an error, triggering the data synchronization mechanism.

[0205] Subsequently, once the difference between the displayed content of the two communication scheduling screens exceeds a difference threshold, the communication system triggers an automatic correction process based on a data synchronization mechanism. For example, this can prioritize synchronization and updating of critical data based on a preset priority, such as the timeliness of the alarm information. Triggering the automatic correction process based on the data synchronization mechanism, for example, includes sending a synchronization instruction to the two communication scheduling screens whose displayed content differs by more than the difference threshold.

[0206] Error capture can be integrated into the prompt template and call the visual language model to synchronize the displayed content.

[0207] As an implementation method, the data synchronization mechanism can be triggered through a rule-based engine. In other words, a rule-based engine is designed to dynamically select and apply the appropriate data synchronization mechanism, thereby matching the error size with the data synchronization mechanism.

[0208] Based on this, the communication system can use the visual language model to periodically or in real time perform difference detection on multiple communication scheduling screens, and when the difference detection indicates that the display content difference between at least two communication scheduling screens is greater than the difference threshold, the difference value is matched with the predefined rules and the corresponding data synchronization mechanism is executed, such as one or more of data retransmission, image rerendering, regeneration of layout data, and issuing an alarm.

[0209] For example, in scenario 1, the corresponding prompt content may be set as:

[0210] Imagine you're a dynamic error detection and content synchronization system. The tasks you perform include: detecting that screen A displays fiber optic cable fault data, screen B displays alarm status information, and that the two screens differ in content. Specifically, screen A displays a chart of the fiber optic cable fault situation; screen B displays the current alarm status of the fiber optic cable. Current error detection is as follows: Time lag detection: The timestamp T1 on screen A differs by several hours from the timestamp T2 on screen B. The calculated time difference Δt is T1-T2. Numerical difference detection: The fiber core utilization rate on screen A is A, the number of alarms on screen B is B, and the numerical difference is |AB|. Image difference detection: The snapshot SSIM value of screens A and B is 0.7 (ideally close to 1), indicating significant differences in displayed content.

[0211] The output format is: Please generate synchronization instructions to ensure the information on screens A and B is synchronized. The specific synchronization strategy is as follows: 1. Predicting the user's next action based on real-time prediction technology: By analyzing the historical operation patterns of operators A and B, we can predict their likely next steps. For example, we predict that operator A may further analyze the impact area of a fiber optic cable fault, while operator B may review historical alarm trends. Synchronize the two screens in advance to ensure consistent information before the next action, minimizing misunderstandings caused by inconsistent data. 2. Based on data association and synchronization technology: The communication system must ensure the data correlation between screens A and B. If the time difference Δt exceeds a set threshold, the data synchronization mechanism is immediately triggered. If the difference between the number of faults A and the number of alarms B exceeds a threshold, the system automatically synchronizes the alarm status and fault information to ensure data consistency between the two screens. 3. Knowledge-graph-based question-answering and analysis: When operators encounter inconsistencies between fault and alarm data, knowledge graph technology is used to automatically analyze the correlation between the two and provide suggestions for correlating alarms and faults. For example, the system can infer the specific cause of an alarm caused by a fiber optic cable fault based on the knowledge graph and display the relevant information simultaneously on screens A and B. 4. Based on multimodal data fusion technology: The system performs multimodal data fusion on fault charts and alarm status information. By integrating the charts and alarm status into a unified, comprehensive view, it is displayed simultaneously on both screens A and B, allowing operators to obtain a complete early warning situation from both screens. 5. Synchronization strategy: Time synchronization: When the time difference Δt exceeds the set threshold, the system immediately synchronizes the alarm status on screen B with the fault data on screen A to ensure consistent timestamps. Data synchronization: When the numerical difference between the number of faults A and the number of alarms B exceeds the threshold, the system automatically synchronizes the alarm information to ensure that the data on screens B and A are consistent. Image synchronization: When the image difference SSIM value falls below the set threshold, the system automatically updates the charts and alarm status on screens A and B to ensure real-time synchronization of image content.

[0212] In addition, in some embodiments, in order to prevent errors or "hallucinations" encountered when the visual language model updates the display interface, that is, the model generation results do not match the actual situation, the communication system can support determining the user's attention information based on the user's interaction behavior with the communication scheduling screen. Among them, the interactive behavior includes operational interaction and dialogue interaction. Exemplarily, the operational interaction refers to monitoring the user's operations on the communication scheduling screen, including actions such as mouse hovering, clicking, dragging, and zooming. For example, if a user clicks on or views a specific alarm area, the communication system will regard it as the user's current attention information. The dialogue interaction refers to the dialogue between the user and the communication scheduling screen (communication system). For example, if the user mentions "the progress of the processing of a certain optical cable alarm" many times in the conversation, the communication system will identify it as the user's attention information.

[0213] For the two users corresponding to the two communication scheduling screens indicated by the multi-screen display content monitoring task, if the attention information of the two users is different, and the attention information of at least one user is not displayed on the communication scheduling screen corresponding to the other user, the synchronization mechanism is triggered to synchronously display the attention information of the two users on the communication scheduling screen corresponding to the other user, thereby ensuring that the display of the two communication scheduling screens is consistent and improving the accuracy of the layout.

[0214] For example, in a scenario where at least two communication scheduling screens are used to assist in task processing, if the two users of the two communication scheduling screens are User A and User B, and monitoring detects that User A repeatedly checks the historical fault records of the optical cable, while the topic of the conversation between User B and the communication system is regional alarm processing, it can be determined that User A and User B's information of interest is different. Based on this, the displayed content on User A's communication scheduling screen can be further matched with User B's information of interest to obtain a first matching result, and the displayed content on User B's communication scheduling screen can be matched with User A's information of interest to obtain a second matching result. If the first matching result indicates a mismatch, it means that the information User B is interested in is not displayed on User A's communication scheduling screen. Similarly, if the second matching result indicates a mismatch, it means that the information User A is interested in is not displayed on User B's communication scheduling screen. This is clearly not conducive to task assistance. Therefore, a visual language model can be used to obtain the HTML content of the communication scheduling screens corresponding to the two users, extract the data display areas that need to be synchronized, and extract the content difference data between the two users' communication scheduling screens. This content difference data can then be injected into the corresponding HTML content of the two communication scheduling screens to ensure consistent display of the two communication scheduling screens.

[0215] In addition, after receiving the synchronization instruction, the communication scheduling screen can consume, read and display data from the same message queue. For example, a distributed database or message queue can be used to achieve real-time synchronization of data between screens. Exemplarily, a distributed message queue system such as Kafka or RabbitMQ can be used. Specifically, a message queue can be set between two communication scheduling screens, and whenever warning data is generated, the warning data will be pushed to the message queue. As a consumer of the message queue, the communication scheduling screen can read the warning data from the message queue and display it synchronously. After the consumer confirms the message, the message queue will mark the data as processed to ensure that the data is not lost, thereby achieving data synchronization.

[0216] In addition, you can also perform synchronous monitoring of the communication scheduling screen. Synchronous monitoring here refers to the monitoring performed when the communication scheduling screen synchronizes the display content in response to synchronization instructions. For example, Prometheus combined with Grafana can be used for monitoring.

[0217] In addition, in order to improve the synchronization efficiency of displayed content, a low-latency communication protocol can also be used for delay control. The low-latency communication protocol includes, for example, WebSocket or gRPC (high-performance remote procedure call protocol).

[0218] Continuing with scenario 1, when fiber optic cable fault information or alarm information is updated, the communication system can push the updated data to a message queue using tools such as Kafka. The two communication scheduling screens to be synchronized act as consumers, fetching the latest data from the queue and displaying it. After each communication scheduling screen receives and displays the data, it sends a confirmation message to the message queue to ensure correct data processing and prevent data loss. Furthermore, data synchronization between the two communication scheduling screens can be monitored using tools such as Prometheus and Grafana, collecting and displaying latency and data consistency metrics in real time. If data delays or synchronization inconsistencies are detected between the two communication scheduling screens, an alert is immediately issued, notifying operations and maintenance personnel to address the situation. Furthermore, a bidirectional communication link can be established between the two communication scheduling screens using WebSocket to ensure real-time data transmission. When the fault chart on one communication scheduling screen is updated, the alarm information on the other is also updated synchronously, ensuring that users receive consistent information on different screens.

[0219] It can be seen that this embodiment couples multi-screen information synchronization technology with the prompt word engineering of the visual language model. When a significant error is found between the two screens, a decision is made automatically to reduce human intervention. At the same time, it can also ensure that the information between the two communication scheduling screens remains consistent and complementary. The two screens can accurately and synchronously process and display warning information, and can dynamically adjust the display content to ensure that the information on the two screens complement each other, which helps team members better understand the overall situation and avoid communication barriers caused by inconsistent information. At the same time, when the information on one screen changes, the other screen can automatically make corresponding adjustments and guide the user's attention through intelligent prompts or highlights. At the same time, it also ensures that the data between the two screens is synchronized in real time, ensuring consistency and accuracy during information transmission, so that team members can see updates immediately, effectively preventing display errors caused by communication delays or errors, and improving the efficiency of collaborative work.

[0220] Based on the large-scale model agent-based communication scheduling and large-screen monitoring method described in any of the above embodiments, this application also provides a computer program product, which includes one or more computer programs or instructions. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium. When the computer program is executed by a processor, it implements the large-scale model agent-based communication scheduling and large-screen monitoring method described in any of the above embodiments.

[0221] Based on the communication scheduling large-screen monitoring method based on a large model agent described in any of the above embodiments, the present application also provides a communication scheduling large-screen monitoring system based on a large model agent, wherein the large model agent includes a visual language model, and the communication scheduling screen is used to display monitoring data, environmental data, fault data, and emergency data of the communication optical cable. The system includes:

[0222] A training module for performing supervised training on a link prediction model using a historical question dataset; wherein the historical question dataset includes multiple historical question links carrying features, each of which includes a historical question content and subsequent question content of a user regarding the display content of a communication scheduling screen; the features include text features and context features of the historical question links;

[0223] a prediction module, configured to obtain the current question content of the user regarding the communication optical cable, and to obtain a predicted question link of the current question content using a trained link prediction model, wherein the predicted question link includes at least the next question content regarding the communication optical cable;

[0224] a layout module, configured to generate prompt content based on the predicted question link, current display data of the communication scheduling screen, and a preset prompt template, and input the prompt content into a trained visual language model to obtain layout data generated by the visual language model; wherein the layout data includes data to be displayed and display parameters; the data to be displayed is related to the predicted question link and includes one or more of the monitoring data, the environmental data, the fault data, and the emergency data;

[0225] An updating module is configured to update a display interface of the communication scheduling screen based on the layout data.

[0226] Based on the communication scheduling large-screen monitoring method based on a large model intelligent agent described in any of the above embodiments, this application also provides Figure 3 A schematic diagram of the structure of an electronic device is shown in FIG. Figure 3 At the hardware level, the electronic device includes a processor, an internal bus, a network interface, memory, and non-volatile storage, and may also include other hardware required for its services. The processor reads the corresponding computer program from the non-volatile storage into the internal memory and then runs it to implement the large-scale model agent-based communication scheduling large-screen monitoring method described in any of the above embodiments.

[0227] The present application also provides a computer storage medium storing a computer program. When the computer program is executed by a processor, it can be used to execute a communication scheduling screen layout method based on a visual language model as described in any of the above embodiments.

[0228] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.

[0229] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0230] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0231] The foregoing is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.

[0232] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

[0233] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

Claims

1. A communication scheduling large-screen monitoring method based on a large model agent, characterized in that: The large model agent includes a visual language model, and the communication scheduling screen is used to display monitoring data, environmental data, fault data, and emergency data of the communication optical cable. The method includes: A link prediction model is supervisedly trained using a historical question dataset; wherein the historical question dataset includes multiple historical question links carrying features, each of which includes a historical question content and subsequent question content of a user regarding the display content of a communication scheduling screen; the features include text features and context features of the historical question links; Obtaining a current question content of a user regarding the communication optical cable, and using a trained link prediction model to obtain a predicted question link of the current question content, wherein the predicted question link includes at least the next question content regarding the communication optical cable; The visual language model is trained using a difference dataset of screen display content; the difference dataset includes time lag type data, numerical difference type data, and image difference type data. Each type of difference data contains screen data of two communication scheduling screens and processing operations on the two communication scheduling screens. The screen data includes a timestamp, an alarm number, and an image snapshot. The processing operations include synchronization operations and data update operations. Prompt content is generated based on the predicted question link, current display data of the communication scheduling screen, and a preset prompt template, and the prompt content is input into a trained visual language model to obtain layout data generated by the visual language model; wherein the layout data includes data to be displayed and display parameters; the data to be displayed is related to the predicted question link and includes one or more of the monitoring data, the environmental data, the fault data, and the emergency data; updating a display interface of the communication scheduling screen based on the layout data; Using a visual language model, based on the multi-screen display content monitoring task in the prompt content, difference detection is performed on the display content of at least two communication scheduling screens located in different locations and used for collaborative management of communication optical cables; the difference detection includes time synchronization difference detection, monitoring value difference detection, and image content difference detection; If the difference detection indicates that the difference in displayed content is greater than a difference threshold, outputting a synchronization instruction to at least two communication scheduling screens so that the communication scheduling screens consume, read, and display data from the same message queue; The user's attention information is determined based on the interaction behavior of each user with the communication scheduling screen. If the attention information of the users corresponding to at least two communication scheduling screens is different, and the attention information of at least one user is not displayed on the communication scheduling screen corresponding to the other user, the attention information of the two users will be synchronously displayed on the communication scheduling screen corresponding to the other user.

2. The method according to claim 1, characterized in that The prompt template also includes associations between multiple data; wherein the associations are pre-existingly obtained by mining multiple rounds of interaction data between the user and the communication scheduling screen using association rules; the associations include an association between the monitoring data of the communication optical cable and the environmental data, an association between the fault data and the emergency data, and an association between the monitoring data of the communication optical cable and the fault data; The generated prompt content carries the association relationship, so that the layout data generated by the visual language model includes first data to be displayed related to the predicted question link, and second data to be displayed related to the first data to be displayed based on the association relationship; wherein the association between the first data to be displayed and the second data to be displayed is represented in the form of a chart and / or a map.

3. The method according to claim 1, characterized in that The prompt content further includes one or more prompt information, so that the layout data generated by the visual language model also includes third display data obtained based on the prompt information; the prompt information includes: Prompt information for prompting data related to the current question content mined from a preset database, wherein the third display data includes the data related to the current question content obtained by mining; Prompt information for prompting the user to call a multimodal pre-trained model to obtain fused data, wherein the third display data includes fused data obtained by mapping data related to the current question content obtained from multiple data sources into the same vector space using the multimodal pre-trained model; wherein the multimodal pre-trained model is trained using text, image, and video data historically displayed on the communication scheduling screen; Prompt information for invoking a knowledge graph, wherein the third display data includes knowledge data in the knowledge graph related to the current question content; wherein the knowledge graph is constructed using relationships between multiple entities of the optical cable; The prompt information is used to prompt the calling of the adaptive model, and the third display data includes display data obtained by using the adaptive model based on the context and / or scene of the current question content.

4. The method according to claim 1, wherein The prompt content also includes the current scene information input by the user; the data to be displayed is related to the current scene information; wherein, If the current scenario information is a daily communication scheduling scenario, the data to be displayed includes: the monitoring data, the environmental data, and the fault data; If the current scene information is an emergency command scene, the data to be displayed includes: optical cable topology data of the emergency area, alarm data in the emergency area, and security data of the emergency area; If the current scene information is a holiday-themed scene, the data to be displayed includes: monitoring data of the target communication optical cable in the holiday protection area, alarm information in the holiday protection area, and management information of the holiday protection area; If the current scenario information is a reporting scenario, the data to be displayed includes: indicator statistical data, indicator analysis data, and service guarantee status data of the communication optical cable.

5. The method according to claim 4, characterized in that The visual language model is trained using screenshots carrying target information, wherein the target information includes: question-answer pairs about the screenshots, logical relationship information between elements in the screenshots, navigation operation information, and summary information; The prompt content includes the current scene information and user instructions input by the user, so that the visual language model executes the user instructions on the data to be displayed corresponding to the current scene information; the user instructions include one or more of screen information understanding instructions, question and answer instructions, navigation instructions and summary instructions.

6. The method according to claim 1, characterized in that The prompt content also includes the current question content and question-structured task information, and the method further includes: Converting the current question content into a structured question and generating an answer content based on the question structuring task information through the visual language model; The question structuring task includes determining the scope of the question topic, determining the type of data involved in the question, building a question framework, refining the question elements, and setting the answer format; Determining the scope of the question topic includes determining whether the current question content is centered around communication optical cable related data displayed on the communication scheduling screen; Determining the data type involved in the question includes categorizing and organizing the data type corresponding to the current question content, the data type including optical cable core utilization rate, environmental parameters, impact of environmental parameters on optical cable performance, computer room temperature thresholds and countermeasures, faulty optical cable information, and unresolved optical cable faults; The constructing of the question framework includes converting the current question content into one or more of general questions, specific numerical questions, comparison questions, cause analysis questions, impact assessment questions, prediction questions, and solution questions based on different data categories; The refinement of question elements includes refining the time, space, object, condition, and cause elements involved in the current question content; The setting of the answer format includes specifying the answer format, and the answer format includes a list, a chart, and text.

7. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, when the processor calls the executable instruction, the method according to any one of claims 1 to 6 is implemented.

8. A communication scheduling large-screen monitoring system based on a large model intelligent agent, characterized in that: The large model agent includes a visual language model, and the communication scheduling screen is used to display monitoring data, environmental data, fault data, and emergency data of the communication optical cable. The system includes: A training module for performing supervised training on a link prediction model using a historical question dataset; wherein the historical question dataset includes multiple historical question links carrying features, each of which includes a historical question content and subsequent question content of a user regarding the display content of a communication scheduling screen; the features include text features and context features of the historical question links; a prediction module, configured to obtain the current question content of the user regarding the communication optical cable, and to obtain a predicted question link of the current question content using a trained link prediction model, wherein the predicted question link includes at least the next question content regarding the communication optical cable; a layout module, configured to generate prompt content based on the predicted question link, current display data of the communication scheduling screen, and a preset prompt template, and input the prompt content into a trained visual language model to obtain layout data generated by the visual language model; wherein the layout data includes data to be displayed and display parameters; the data to be displayed is related to the predicted question link and includes one or more of the monitoring data, the environmental data, the fault data, and the emergency data; An updating module is configured to update a display interface of the communication scheduling screen based on the layout data.

9. A computer-readable storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed by a processor, the steps of any one of the methods of claims 1-6 are implemented.

Citation Information

Patent Citations

  • Data processing method and device and electronic equipment

    CN112148939A