A voice dispatching management and control method and system for mine multi-service linkage
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANKE HIGH TECH (HENAN) RES INST CO LTD
- Filing Date
- 2026-06-12
- Publication Date
- 2026-08-07
AI Technical Summary
[0005]为了解决现有矿山语音调度管控方法无法对复合语音指令中的业务意图和呈现意图进行协同理解,且无法根据语义内容自动确定对应的大屏展示方式和布局策略的问题,本发明提供了一种面向矿山多业务联动的语音调度管控方法,所述方法包括:
本发明通过构建复合意图知识图谱、联合意图理解模型以及智能呈现推理机制,实现了业务操作意图与屏幕呈现意图的统一理解和协同决策,使系统不仅能够理解用户想查询什么,还能够理解用户希望如何展示,从而实现语音指令到业务处理、大屏布局以及画面呈现的自动闭环联动,显著提升矿山调度场景下的人机交互智能化水平。
Smart Images

Figure CN122531389A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent scheduling and voice human-computer interaction technology in mines, specifically to a voice scheduling and control method and system for multi-business linkage in mines. Background Technology
[0002] The mine dispatch center is the command hub for coal mine safety production. Dispatchers need to simultaneously monitor real-time data from multiple business systems, including personnel positioning, safety monitoring, video surveillance, and production automation, and make rapid decisions based on the situation on-site. In recent years, voice interaction technology has been gradually introduced into industrial dispatching scenarios to replace some manual operations and reduce the workload of dispatchers. A typical implementation of existing mine voice control solutions involves pre-defining several fixed voice command templates, each bound to a specific system function or screen switching action. The system uses a voice recognition engine to transcribe the operator's voice into text, matches it against the command template library, and triggers the corresponding single operation upon successful matching, such as opening a monitoring screen or querying specific data.
[0003] However, existing mine voice dispatching schemes have the following main drawbacks: First, the interaction mode is simplistic and rigid, unable to handle complex intents. Existing solutions only support a mapping mode where one command corresponds to one single operation. In actual scheduling scenarios, operators' voice commands often contain complex intents. For example, requesting to view the gas concentration data from the second face of a fully mechanized mining operation simultaneously includes the business operation intent of querying gas concentration data and the intent of displaying the image on a large screen. Existing solutions require breaking down such complex commands into multiple single commands for separate execution, which is cumbersome and does not conform to natural human expression habits, failing to truly achieve intelligent interaction where what you say is what you see.
[0004] Second, there is a lack of intelligent connection between voice commands and large-screen presentation, and the presentation method is highly dependent on manual specification. In the existing solution, voice commands are only responsible for triggering a predefined data query or screen retrieval. As for how the query results should be displayed in which area of the large screen in which form (list, curve, map, video), it depends entirely on the operator's prior configuration in the background or manual adjustment. Summary of the Invention
[0005] To address the shortcomings of existing mine voice dispatch and control methods, which fail to collaboratively understand the business and presentation intentions within complex voice commands and cannot automatically determine the corresponding large-screen display method and layout strategy based on semantic content, this invention provides a voice dispatch and control method for multi-business linkage in mines. The method includes: S1. Construct a composite intent knowledge graph in the mining field, and obtain a joint intent understanding model based on the composite intent knowledge graph; S2. Based on the pre-trained mine background noise suppression model, a specific wake-up word is obtained to switch the voice acquisition device from sleep mode to working mode, and the voiceprint features of the current operator are obtained to enable directional sound reception; S3. Convert the collected speech into a text sequence, and input the text sequence into the joint intent understanding model to obtain a scheduling instruction set; S4. Execute the scheduling instruction set to acquire multi-source heterogeneous business data and drive the dynamic reassembly of the physical large screen image; send data query requests in parallel to system interfaces such as personnel positioning and security monitoring; simultaneously, send screen layout instructions to the large screen splicing processor. The large screen image immediately begins to respond, such as clearing the screen first and then establishing a connection to the specified window. This step achieves the visual preparation of "what you see is what you get".
[0006] S5. Render the multi-source heterogeneous business data onto the physical large screen and generate voice broadcast feedback simultaneously. After receiving the JSON / XML data returned by each business system, render the data in real time to the corresponding video window according to the presentation method specified in the screen layout instructions (such as overlaying personnel list data onto the GIS map, and drawing the gas concentration value-driven curve control).
[0007] Synchronous voice broadcast: At the same time, the key summary information of the query (such as the current total number of people underground is 328, and everything is normal) is broadcast through voice synthesis, and the virtual human can be driven to lip-sync, completing the closed loop of human-computer dialogue.
[0008] This method constructs a joint intent understanding model to simultaneously parse multiple business and display semantics within a single voice command. This enables the mine scheduling system to directly understand complex intent expressions in natural language, reducing manual command splitting, improving voice scheduling efficiency and naturalness of interaction, and achieving a WYSIWYG intelligent scheduling interaction mode. By constructing an inference rule base and a large-screen digital twin model, the system can automatically generate the optimal display strategy based on the data attributes, entity types, spatial relationships, and user job characteristics of the query object. This achieves intelligent mapping between business data and large-screen presentation, reducing reliance on manual configuration and improving large-screen display efficiency and scheduling decision-making efficiency.
[0009] Furthermore, the specific steps of S1 include: S101: Obtain all mining business entities and establish a colloquial thesaurus for each of the mining business entities; Construct queryable attributes for each of the mine-wide business entities, map the queryable attributes to the query interface of the backend business system, and construct triples based on the mine-wide business entities, the queryable attributes, and the query interface; Construct conditional slots for the queryable attributes, wherein the conditional slots include slot type, colloquial trigger word library and normalization transformation rules, and the slot type includes time range, spatial region, entity status, threshold comparison and identity identifier; Based on the intent of generating conditional combinations for the aforementioned conditional slots; Based on the combined conditions, filter parameters for the query interface are generated. Based on the triples, a semantic association layer of entity-attribute-query interface is constructed. S102: The physical screen is virtualized into a computable canvas, and the logical areas of the canvas and the area attributes of the logical areas are constructed. A preset visual control is constructed, and the control properties of the visual control are constructed. The control properties include the data source types that can be bound and the minimum set of parameters required.
[0010] Based on the logical region, the visual control, and the control-region adaptation rules, a dynamic screen layout template is constructed. Based on the aforementioned dynamic screen layout template, a knowledge layer for presenting large screen logical areas and visual controls is constructed. S103: Construct composite intent templates for different mine scheduling scenarios, each composite intent template containing a query intent and a presentation intent; Based on the semantic association layer, the presentation knowledge layer, and the composite intent template, the composite intent knowledge graph is constructed.
[0011] A composite intent knowledge graph is constructed for the mining industry, unifying and associating mining entities, entity attributes, query interfaces, screen layouts, control presentation methods, and scheduling scenario templates. After obtaining text commands through speech recognition, a joint intent understanding model constructed using the composite intent knowledge graph is used to achieve: speech-structured intent-scheduling command-data retrieval-dynamic large screen presentation. This completes the integrated linkage between business and display, enabling the system to directly convert speech commands into data query and large screen display actions, significantly reducing the complexity of scheduling operations and improving the efficiency of multi-business linkage in the mining industry.
[0012] Furthermore, before executing S4, the method further includes: Construct an occupancy status table for the logical area, which includes the currently occupied display task, the current priority level of the task, the timestamp when the task started occupying the area, whether the area can be preempted, and the core business entity associated with the currently displayed content. Obtain the logical relationships between the logical regions to obtain a cross-regional relationship diagram; Build a global history stack; Based on the occupancy status table, the cross-regional relationship diagram, and the global display history stack, a multi-dimensional digital twin model of the physical large screen display status is constructed. The scheduling instruction set is input into the multidimensional digital twin model to obtain the conflict results; The conflict results are matched with the conflict resolution rule base to obtain the target resolution rule, and an optimized scheduling instruction set is obtained based on the target resolution rule. The conflict resolution rule base includes resolution strategy rules under different conflict scenarios.
[0013] Existing voice dispatch and control methods in mines lack real-time perception and adaptive adjustment capabilities for screen display status, leading to frequent screen conflicts in multi-tasking scenarios. Current solutions typically employ overlay or fixed-area screen projection strategies when receiving new voice commands. When the large screen is highlighting alarm content or monitoring scenarios of concern to superiors, new ordinary query commands may forcibly overwrite critical content; or when multiple areas are already occupied, the new screen cannot find a suitable display location. The system lacks a comprehensive understanding of the current screen occupancy status, task priorities, and content relevance, and cannot flexibly coordinate the layout of multi-tasking screens like an experienced dispatcher, resulting in a clunky interactive experience and potential security risks.
[0014] This method constructs an occupancy status table, a cross-regional relationship graph, and a global display history stack, further forming a digital twin model of the physical large screen display status. Before executing a new instruction, it first detects regional conflicts, priority conflicts, and preemption conflicts in the digital twin model, thereby achieving prediction, realizing real-time mapping and pre-simulation of the display status. Potential conflicts can be detected before the scheduling instruction is executed, reducing the accidental coverage of critical monitoring screens, improving the utilization rate of large screen resources and scheduling security.
[0015] Furthermore, the resolution strategy rules include: Rule 1, absolute priority preemption: emergency preemption, alarm-type tasks have the highest priority, directly preempting the target area of non-alarm-type tasks; Rule 2, Content Collaboration and Coexistence: In-situ attribute switching, without creating a new window, directly switches the data source attributes of the controls bound to the current area, or overlays a temporary floating attribute panel on top of it; Rule 3, Intelligent Space Allocation: Flexible degradation allocation, automatically assigning new screens to the available candidate area with the highest matching value with the content type, and adaptively adjusting the window size to the optimal; Rule 4, Information Density Overload: Temporary Overlay and Fade-out. New content is overlaid on the existing content with the highest relevance in the form of a semi-transparent floating window, and automatically fades out after a preset time without operation. Rule 5, Context-Aware Clearing: Associated cascading clearing and focusing. When clearing an area, it automatically identifies and clears auxiliary screen areas that are business-related to that area.
[0016] Five conflict resolution strategies are established for conflict scenarios. The corresponding processing method is automatically selected according to different conflict types. Compared with the fixed window management mechanism, this embodiment can dynamically select the conflict resolution strategy according to business priority, entity relationship and display status, so that the large screen display has stronger adaptability and business continuity, and reduces information chaos caused by frequent window opening and closing.
[0017] Furthermore, the specific steps for matching the conflict results with the conflict resolution rule base to obtain the target resolution rule include: Based on the conflict results, a conflict feature vector is constructed, which includes request features, current occupancy features, and global situation features. Each of the aforementioned resolution strategy rules is converted into a conditional reasoning decision tree, and each rule node of the conditional reasoning decision tree contains a conditional expression and an action instruction; Based on the priority level, the conflict feature vector and the conditional reasoning decision tree are matched to obtain the target resolution rule.
[0018] The conflict scenario is transformed into a conflict feature vector, including request features, current occupancy features, and global situation features. At the same time, the rule base is transformed into a conditional reasoning decision tree. The feature vector is used to match the decision tree to achieve automatic rule selection, reduce manual traversal of conflict rules, improve rule matching efficiency, enable the system to quickly determine the optimal conflict resolution strategy based on the real-time operating status, and enhance the system's intelligence level and real-time response capability.
[0019] Furthermore, if there are two or more resolution strategy rules with the same priority level and whose conditional expressions all satisfy the conditions, the method further includes: Based on the resolution strategy rules, candidate rules are obtained, action instructions for the candidate rules are acquired, and a candidate operation list is obtained. A real-time state snapshot is obtained based on the aforementioned multidimensional digital twin model; Input the candidate operation list into the screen digital twin sandbox to obtain a snapshot of the predicted state; The operation cost is obtained by comparing the real-time state snapshot with the predicted state snapshot. Obtain the resolution strategy rule corresponding to the minimum value of the operation cost, and obtain the target resolution rule; The operational costs include visual stability costs, cognitive continuity costs, and information relevance costs. The visual stability cost includes window creation cost, window closing cost, and window movement / scaling cost; The cognitive continuity costs include entity-control unbinding costs, attribute-control in-situ switching costs, and data flow interruption and continuation costs. The cost of information correlation includes the cost of creating information silos and the cost of destroying information clusters.
[0020] To address situations where multiple rules simultaneously meet certain conditions, a system is constructed that includes real-time state snapshots, a screen digital twin sandbox, and predicted state snapshots. Visual stability costs, cognitive continuity costs, and information relevance costs are designed to form a comprehensive operational cost function. Different strategy execution results are simulated within the digital twin sandbox, and the solution with the minimum total cost is selected. A digital twin sandbox prediction mechanism is introduced, enabling the quantitative selection of conflict resolution strategies through multi-dimensional cost evaluation. This not only reduces the frequency of interface transitions but also maintains the scheduler's cognitive continuity and the integrity of information relevance, improving human-machine collaboration efficiency in complex scenarios.
[0021] Furthermore, the specific steps of S3 include: The text sequence is input into a joint intent network, which includes a shared encoder and parallel business operation intent decoders and screen presentation intent decoders, and the business operation intent decoders and screen presentation intent decoders are connected by cross attention. Based on the shared encoder, the text sequence is parsed to obtain several minimum input units and the contextual semantic representation of the minimum input units; Based on the business operation intent decoder and screen presentation intent decoder, the smallest input unit is parsed to obtain intent data and presentation data; The intent data and the presentation data are mapped to the joint intent understanding model to obtain the target intent template; the scheduling instruction set is obtained based on the target intent template.
[0022] A shared encoder, a business operation intent decoder, and a screen presentation intent decoder are constructed, and the two decoders work collaboratively through cross-attention. This achieves unified parsing of a single voice command containing both business and display intents. Simultaneously, it parses both the query query and the display requirements, addressing the problem that traditional voice systems can only recognize business operations but not display needs. This enables synchronous parsing of business and display semantics, improving the accuracy of complex scheduling command recognition and the consistency of display results.
[0023] Furthermore, the specific steps to obtain the target intent template include: Based on the colloquial trigger word library, the intent data, and the presentation data, a candidate entity list is generated; The candidate entity list and the composite intent knowledge graph are matched to obtain the candidate entity graph; Obtain the semantic similarity between the context semantic representation and the candidate entity graph to obtain the target data; The target data is mapped to the joint intent understanding model to obtain the target intent template.
[0024] Construct a candidate entity list and a candidate entity graph subgraph, calculate the matching degree between the context semantics and the candidate entity graph, realize entity disambiguation, solve the ambiguity problem caused by a large number of abbreviations, common names and entities with the same name in the mining field, improve the accuracy of entity recognition, and ensure that subsequent query and scheduling actions are applied to the correct business objects.
[0025] Furthermore, if there are information gaps in the target intent template, the step of obtaining the scheduling instruction set based on the target intent template further includes: Construct a reasoning rule base, which includes reasoning rules for different reasoning types, including data type-control adaptation reasoning, entity type-default area reasoning, spatial semantics-associative presentation reasoning, role / position-view preference reasoning, and status-emergency presentation reasoning. Based on the inference rule base, the target intent template, and the target data, the missing information is obtained; Based on the missing information and the target intent template, the scheduling instruction set is obtained.
[0026] Existing voice dispatch and control methods in mines suffer from a lack of intent completion capabilities for ambiguous commands and place excessively high demands on the precision of operator expression. Current solutions rely on operators issuing commands that precisely match preset templates. When operators use abbreviations, pronouns (such as "its monitoring," "the previous screen"), or fail to specify a display method, the system cannot utilize the mine's business knowledge graph and dialogue context for reasoning and completion, leading to interaction failures or requiring multiple rounds of clarification, thus reducing dispatch efficiency.
[0027] By building a reasoning rule base, when the target intent template is missing parameters, the system can automatically fill in the missing information through reasoning rules. Users do not need to express all parameters completely. The system can automatically fill in the query object, display area and related content based on knowledge graph and rule reasoning, thereby improving the error tolerance of voice interaction and natural language understanding capabilities.
[0028] This invention provides a voice dispatch and control system for multi-service linkage in mines, the system comprising: Graph Unit: Used to construct a composite intent knowledge graph in the mining field, and to obtain a joint intent understanding model based on the composite intent knowledge graph; Start-up unit: Based on the pre-trained mine background noise suppression model, it acquires a specific wake-up word, switches the voice acquisition device from sleep mode to working mode, acquires the voiceprint features of the current operator, and starts directional sound reception; Instruction unit: used to convert the collected speech into a text sequence, input the text sequence into the joint intent understanding model, and obtain a scheduling instruction set; Execution unit: used to execute the scheduling instruction set, acquire multi-source heterogeneous service data, and drive the dynamic reorganization of the physical large screen screen; Feedback unit: used to render the multi-source heterogeneous business data to the physical screen and generate voice broadcast feedback simultaneously.
[0029] The principle and beneficial effects of this system are similar to those of this method, and will not be elaborated further on here.
[0030] One or more technical solutions provided by this invention have at least the following technical effects or advantages: This invention achieves unified understanding and collaborative decision-making between business operation intent and screen presentation intent by constructing a composite intent knowledge graph, a joint intent understanding model, and an intelligent presentation reasoning mechanism. This enables the system to not only understand what the user wants to query, but also how the user wants to display the information. As a result, it realizes automatic closed-loop linkage from voice commands to business processing, large screen layout, and screen presentation, significantly improving the level of human-computer interaction intelligence in mine scheduling scenarios. Attached Figure Description
[0031] The accompanying drawings, which are provided to further illustrate embodiments of the invention and constitute a part of this invention, are not intended to limit the scope of the invention. Figure 1 This is a flowchart illustrating a voice dispatch and control method for multi-service linkage in mines, as described in this invention. Detailed Implementation
[0032] To better understand the above-mentioned objectives, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, where there is no conflict, the embodiments of the present invention and the features thereof can be combined with each other.
[0033] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.
[0034] Example 1 refer to Figure 1This embodiment provides a voice dispatch and control method for multi-service linkage in mines, the method including: S1. Construct a composite intent knowledge graph for the mining sector, and obtain a joint intent understanding model based on the composite intent knowledge graph; use the composite intent knowledge graph to train an existing neural network model or deep learning model to obtain the joint intent understanding model. The specific steps of S1 include: S101: Acquire all mine business entities (such as personnel, sensors, cameras, and working faces), and establish a colloquial thesaurus for each of the mine business entities; for example, entity: methane sensor, associated with: gas probe, CH4 sensor; entity: fully mechanized mining team 1, associated with: team 1, coal mining shift 1.
[0035] Construct queryable attributes for each of the mine-wide business entities, map the queryable attributes to the query interface of the backend business system, and construct triples based on the mine-wide business entities, the queryable attributes, and the query interface; For example: (Personnel, current location, real-time coordinate query API for personnel positioning system), (Methane sensor, real-time concentration, query API for measurement point value of safety monitoring system), (Camera, video stream address, RTSP stream extraction API for video monitoring platform).
[0036] The condition slots for the queryable attributes are constructed as shown in Table 1. The condition slots include slot type, colloquial trigger word library and normalization conversion rules. The slot type includes time range, spatial region, entity status, threshold comparison and identity identifier. Table 1. Example of a Conditional Slot Based on the intent of generating conditional combinations for the aforementioned conditional slots; For example, a user command: Check the records of gas concentrations exceeding 0.8 in the No. 2 mining area after 8:00 AM today.
[0037] Conditional slot extraction and filling: Spatial area slots: activated and filled by the second mining area, normalized output: spatial area = Area_02; Time range slot: Activated after 8:00 AM today, normalized output: Time range = Start: 2026-01-27 08:00:00, End: Current}; Threshold comparison slot: Activated when the value exceeds 0.8, normalized output threshold comparison = {spatial region: concentration, operator: >, value 0.8}.
[0038] Organize the above condition slots into a structured JSON object as a condition combination intention; Generate the filtering parameters of the query interface based on the combined condition intention, and construct the semantic association layer of entity-attribute-query interface based on the triple. In the triple, find the query interface associated with the historical concentration attribute of the entity gas sensor. Perform parameter mapping and filling on the above combined condition intention, and use it as the filtering parameter of the query interface.
[0039] If the user instruction is: Check where the people in Team 1 are, the combined condition is {identity identifier: Team 1, time range: now}. This intention may trigger a composite call: First, call the personnel organizational structure API according to the identity identifier to query the ID list of all personnel in Team 1. Then, use this personnel ID list as the filtering condition and pass it into the real-time location query API of the personnel positioning system to obtain the current locations of these people at one time.
[0040] S102: Virtualize the physical large screen into a computable canvas, construct the logical areas of the canvas (such as the central main screen area, the left auxiliary screen area, the right auxiliary screen area, the top warning scrolling area, the temporary pop-up window area, etc.) and the area attributes of the logical areas (such as priority, maximum number of accommodated windows, recommended content types). Preset visualization controls (such as GIS map control, real-time list control, trend curve control, video playback control, dashboard control, etc.), and construct the control attributes of the visualization controls. The control attributes include the data source types that can be bound and the minimum parameter sets required. For example, the GIS map control needs to bind longitude and latitude sequence data; the trend curve control needs to bind time-value pair sequences.
[0041] Construct a screen dynamic layout template based on the logical areas, the visualization controls and the control-area adaptation rules. Control-area adaptation rules: Define in which logical areas the controls are suitable to be presented. For example, the video control is suitable for the main screen area or the pop-up window area, and the list control is suitable for the auxiliary screen area.
[0042] Construct the presentation knowledge layer of the large screen logical area-visualization control based on the screen dynamic layout template. S103: Construct composite intention templates for different mine dispatching scenarios. Each composite intention template contains a query intention and a presentation intention. For example: Composite intention template in the personnel patrol scenario: Query intention: {person / team} location - Presentation intention: <GIS map control> locate and highlight in <central main screen area> + <personnel list control> presented in <right auxiliary screen area>.
[0043] Composite intent template for alarm handling scenarios: Query intent: {sensor} Over-limit details - Presentation intent: <trend curve control> presented in the <central main screen area> + associated <video control> automatically popped up in the <temporary pop-up area> + <current value dashboard> presented in the <left auxiliary screen area>.
[0044] Based on the semantic association layer, the presentation knowledge layer, and the composite intent template, the composite intent knowledge graph is constructed.
[0045] S2. Based on the pre-trained mine background noise suppression model, a specific wake-up word is obtained to switch the voice acquisition device from sleep mode to working mode, and the voiceprint features of the current operator are obtained to enable directional sound reception; S3. Convert the collected speech into a text sequence, and input the text sequence into the joint intent understanding model to obtain a scheduling instruction set; S4. Execute the scheduling instruction set, acquire multi-source heterogeneous service data, and drive the physical large screen to dynamically reassemble the screen. S5. Render the multi-source heterogeneous business data onto the physical screen and generate voice broadcast feedback simultaneously.
[0046] Example 2 Based on the above embodiment one, in this embodiment, before executing S4, the method further includes: Construct an occupancy status table for the logical area. The occupancy status table includes the currently occupied display task, the current priority level of the task, the timestamp when the task started occupying the area, whether the area can be preempted (e.g., the emergency alarm screen is set to be non-preemptible), and the core business entity associated with the currently displayed content (e.g., sensor ID, area ID, personnel ID). The logical relationships between the logical regions are obtained to create a cross-regional relationship graph. For example, a region is highlighted on the GIS map in the central main screen area, and the monitoring screen of that region is displayed in the auxiliary screen area on the right. This relationship is recorded as a graph, where nodes are logical regions and edges represent business collaboration display relationships. When the content of the main region is removed, the system can detect that the associated auxiliary regions should also be cleared or updated.
[0047] Build a global display history stack; maintain a display operation history stack triggered by the user via voice or manual operation, supporting scene awareness for undoing and redoing. For example, a user can use voice commands: "Return to the previous screen."
[0048] Based on the occupancy status table, the cross-regional relationship diagram, and the global display history stack, a multi-dimensional digital twin model of the physical large screen display status is constructed.
[0049] The scheduling instruction set is input into the multidimensional digital twin model to obtain the conflict results; For example, the scheduling instruction set is: {Target area: central main screen area, Control: video control, Data: camera 01 stream, Priority: NORMAL}; Query the occupancy status table of the logical area to obtain the current status of the central main screen area; Suppose the query result is: the central main screen area is currently occupied by Task1 with a priority level of CRITICAL, and whether this area can be preempted is false; Conflict determination: A conflict occurs when the priority of the new request's NORMAL is lower than that of the currently occupied CRITICAL, and the current task cannot preempt it.
[0050] The conflict results are matched with the conflict resolution rule base to obtain the target resolution rule, and an optimized scheduling instruction set is obtained based on the target resolution rule. The conflict resolution rule base includes resolution strategy rules under different conflict scenarios. Input the conflict results into the conflict resolution rule base and match them with the target resolution rule.
[0051] The target resolution rules are executed, and based on the multidimensional digital twin model, an idle area that is compatible with the video content type is found, namely the right auxiliary screen area.
[0052] Correction: Change the target area from the central main screen area to the right auxiliary screen area, generate a corrected scheduling instruction set, and calculate the current optimal window size for this area.
[0053] Example 3 Based on the above embodiments, in this embodiment, the resolution strategy rules include: Rule 1, Absolute Priority Preemption: Emergency preemption, alarm-type tasks have the highest priority, directly preempting the target area of non-alarm-type tasks; if the old task to be preempted is suspended to the background task stack, the controls in the preempted area will be unbound from the old data source and bound to the new alarm data source. The state of the suspended task is retained and can be restored with one click after the alarm is cleared.
[0054] Rule 2, Content Collaboration and Coexistence: In-situ attribute switching without creating a new window, directly switching the data source attributes of the controls bound to the current area, or overlaying a temporary floating attribute panel on top of it; for example, if the current central area displays the location of the mining team, and the new instruction is to view their trajectory, the system will identify the same entity, mining team one, and switch the GIS map control from location snapshot mode to trajectory playback mode, instead of opening a new window.
[0055] Rule 3, Intelligent Space Allocation: Flexible degradation allocation, automatically assigning new screens to the available candidate area with the highest matching value with the content type (such as using the cosine similarity algorithm), and adaptively adjusting the window size to the optimal; if the central main screen area is occupied by a permanent task (such as the overall mine production overview), new mining area monitoring screen requests will be automatically assigned to the right auxiliary screen area, and the video window will be enlarged to the recommended size of that area.
[0056] Rule 4, Information Density Overload: Temporary Overlay and Fade-out. New content is overlaid on the existing content with the highest relevance as a semi-transparent floating window, and automatically fades out after a preset time without operation. The system calculates the business relevance of the new content to the content in each current area, and selects the area with the highest relevance (e.g., using a semantic distance algorithm based on a knowledge graph) for overlay. For example, a list of detailed information about Zhang San is overlaid on a GIS map that already displays the location of a team of people, with a voice prompt indicating that the relevant information has been overlaid for you.
[0057] Rule 5, Context-Aware Clearing: Cascading clearing and focusing are associated. When clearing a region, the system automatically identifies and clears auxiliary screen areas that are business-related to that region. When only viewing a specified logical region, the specified logical region is enlarged to the main screen area in full screen, while the rest are closed. The system queries the cross-region relationship graph and marks related nodes as pending clearing. During focusing operations, the target screen controls are dynamically rebound to the central main screen area, and the original central content is downgraded according to Rule 3.
[0058] Example 4 Based on the above embodiments, in this embodiment, the specific steps for matching the conflict result with the conflict resolution rule base to obtain the target resolution rule include: Based on the conflict results, a conflict feature vector is constructed, which includes request features, current occupancy features, and global situation features. The request characteristics include: the priority level of the new instruction (NORMAL), the target logical area expected by the new instruction (central main screen area), the content type expected to be presented by the new instruction (video stream), the core business entity ID associated with the new instruction (camera 2), and the business entity type associated with the new instruction (camera). The current occupancy characteristics include: the priority of the current task in the target area (CRITICAL), whether the target area is currently preemptible (false), the type of content currently displayed in the target area (GIS map), the business entity ID currently associated with the target area (logical area 2), and the business entity type currently associated with the target area (mining area). Global situational characteristics include: the occupancy status vector of all logical areas (e.g., [occupied, idle, occupied, occupied]), whether other areas associated with the new entity have content (true, the associated screen is being displayed in the right auxiliary screen area), and the current operator's role (e.g., safety dispatcher). Each of the aforementioned resolution strategy rules is converted into a conditional reasoning decision tree, and each rule node of the conditional reasoning decision tree contains a conditional expression and an action instruction; Taking rule 3 above as an example: Conditional expression: Condition 1: There is a conflict in the target area; Condition 2: The priority of the new request is insufficient for preemption; Condition 3: Alternative areas exist (excluding currently occupied and non-seizeable areas); Condition 4: Attribute switching scenarios for different entities (excluding the scope of rule 2); Action command: If all the above conditions are met, the target area will be reassigned to the best matching free area, the control window size will be adaptively adjusted according to the new area size, and voice feedback will be generated: [Content summary] has been displayed for you in [Area Name], the screen status table will be updated, and the new area occupancy information will be recorded.
[0059] Based on the priority level, the conflict feature vector and the conditional reasoning decision tree are matched to obtain the target resolution rule.
[0060] Example 5 Based on the above embodiments, in this embodiment, if there are two or more resolution strategy rules with the same priority level and whose conditional expressions all satisfy the conditions, the method further includes: Based on the resolution strategy rules, candidate rules are obtained, action instructions for the candidate rules are acquired, and a candidate operation list is obtained. A real-time state snapshot is obtained based on the aforementioned multidimensional digital twin model; The candidate operation list is input into the screen digital twin sandbox (i.e., virtual large screen simulation environment) to obtain a snapshot of the predicted state; the future state after the application of action commands is simulated. The operation cost is obtained by comparing the real-time state snapshot with the predicted state snapshot. Obtain the resolution strategy rule corresponding to the minimum value of the operation cost, and obtain the target resolution rule; The operational cost includes visual stability cost, cognitive continuity cost, and information relevance cost; the final operational cost is obtained by weighted fusion of these costs. The visual stability cost includes window creation cost, window closing cost, and window movement / scaling cost; this dimension measures the number of new, closed, and moved elements on the screen that will result from executing a sequence of actions, aiming to minimize the creation, destruction, and abrupt changes of elements on the large screen and reduce distractions for the scheduler.
[0061] Window creation cost: Each newly added video window or control instance is considered a higher cost (e.g., weight = 0.5). This is because the appearance of a new window attracts attention and interrupts continuous observation of the existing screen; Window closing cost: Each time an existing window is closed, it is counted as the maximum cost (e.g., weight = 1). Because closing may mean the loss of information, the cost is extremely high if the screen contains critical information that the scheduler has not yet finished viewing. A user gaze duration variable can be maintained. If a window exists continuously for more than 30 seconds, the window closing cost will be multiplied by a cognitive input bonus factor of 1.5. Window movement / scaling cost: Moving a window from one area to another, or changing its size, is considered a medium cost (e.g., weight = 0.2). This can cause temporary visual confusion.
[0062] The visual stability cost is obtained by adding the above costs together.
[0063] The cognitive continuity costs include entity-control unbinding costs, attribute-control in-situ switching costs, and data flow interruption and resumption costs. This dimension measures the degree of disruption in data-entity binding relationships caused by the execution of an action, aiming to maintain the scheduler's cognitive context. This directly supports the continuity of "what is said should be done".
[0064] Entity-control unbinding cost: Replacing the data source of the currently displayed business entity on a control with the data source of a different entity is considered a high cost (e.g., weight = 0.8). For example, switching a window from displaying a team of personnel to displaying a gas concentration curve forcibly resets the dispatcher's cognitive model of that window.
[0065] In-place switching cost of attribute-control: Switching the attribute of the same entity bound to a control (such as from position to trajectory) is considered low-cost (e.g., weight = 0.1). Because the entity context remains consistent, the scheduler's cognitive focus does not need to shift, which is the ideal seamless switching.
[0066] Data stream interruption and resumption cost: If the real-time data stream of a control (such as a gas concentration curve) is shut down, it may create a monitoring blind spot, which will be counted as an additional penalty point (e.g., penalty coefficient = 1.5). The system will check whether the shut-down control is bound to an entity data stream with an alarm threshold association attribute. If so, this penalty will be applied.
[0067] The above costs are added together to obtain the cognitive continuity cost.
[0068] The cost of information relevance includes the cost of creating information silos and the cost of destroying information clusters. This dimension measures the business relevance of new content to existing visual content, encouraging the presentation of related information together to form an information cluster, and conversely penalizing the behavior of scattering related information or destroying existing clusters.
[0069] Cost of creating an information silo: If the entity displayed in the newly created window has no business association (such as spatial association or system association) with any other entity in any other window on the current screen in the knowledge graph, then it is counted as a cost (e.g., weight = 0.3). This is equivalent to creating an information silo on the screen.
[0070] Information Cluster Disruption Cost: If a closing or moving action causes an existing information cluster (i.e., a small group of nodes with high correlation in a cross-regional relationship graph) to disintegrate, it is counted as a cost (e.g., weight = 0.6). For example, the central area displays a map of a mining area, and the secondary screen displays monitoring data for that mining area. These two windows constitute a strongly correlated cluster. If only the monitoring data in the secondary screen is closed, the cost is low; however, if the monitoring data in the secondary screen is forcibly moved to another area unrelated to that mining area, disrupting spatial proximity, this cost will be incurred.
[0071] The above costs are added together to obtain the information correlation cost.
[0072] Example 6 Based on the above embodiments, in this embodiment, the specific steps of S3 include: The text sequence is input into a joint intent network, which includes a shared encoder (such as a pre-trained language model trained on mine production corpus and scheduling instruction logs, such as BERT) and parallel business operation intent decoders and screen presentation intent decoders. The business operation intent decoders and screen presentation intent decoders are connected via cross-attention. When the business operation intent decoder detects an entity type of "camera," it transmits video stream-related semantic biases to the presentation intent decoder through cross-attention, guiding it to decode video controls and pop-up area labels more readily. Conversely, full-screen instructions in the presentation intent also reinforce the semantic selection of detailed data rather than list summaries in the operation intent. This inherent coupling ensures the self-consistency of the composite intent parsing results.
[0073] Based on the shared encoder, the text sequence is parsed to obtain several minimum input units and the contextual semantic representation of the minimum input units, which contains a deep understanding of mine entities, operational verbs, and spatial vocabulary; For example, the instruction would be: Check the gas levels on the second face of the fully mechanized mining operation.
[0074] After analysis, four minimum input units are obtained: [Check, fully mechanized mining face, of, gas].
[0075] A composite intent knowledge graph constructed using S1 is pre-trained to create a model. When a fully mechanized mining face is identified, its vector representation is overlaid with attribute features from the graph, such as entity type (mining area) and coal seam (No. 2 coal seam). The fully mechanized mining face establishes a strong association with gas. The encoder, through attention heads, calculates that this gas is not generic, but rather refers specifically to the gas sensors within the area of the fully mechanized mining face. Ultimately, the gas output vector is no longer merely representations of chemically explosive gas; it is imbued with deeper contextual semantics, including the fully mechanized mining face, the query object, and sensor data.
[0076] Based on the business operation intent decoder and screen presentation intent decoder, the smallest input unit is parsed to obtain intent data and presentation data; Intent data: Action type (e.g., query / control / alarm confirmation / navigation); target business entity and its attributes, identify the boundaries of entities and attributes through sequence labeling, and mark the dependencies between them (e.g., the current location of personnel); filter condition triples, and use relation extraction to identify the (entity, comparison operator, comparison value) structure, such as (gas concentration, >, 0.8%).
[0077] Presentation data: Presentation action type (e.g., throw / switch / zoom in / close / stitch); mention of target presentation area (central main screen / left auxiliary screen, etc.); mention of expected presentation control type (video / curve / map / list); screen operation parameters (e.g., stitching, single screen, rotation).
[0078] The intent data and the presentation data are mapped to the joint intent understanding model to obtain the target intent template; the scheduling instruction set is obtained based on the target intent template.
[0079] Business Operation Intent Decoder: This neural network submodule is responsible for extracting structured information at the operational level from the semantic vectors output by the shared encoder. Its core task is to answer: which business system, which entity, what operation does the scheduler want to perform, and under what conditions?
[0080] Screen presentation intent decoder: A neural network submodule responsible for extracting structured information at the display level. Its core task is to answer: what visualization format the scheduler wants for this information, in which area of the large screen it should be presented, and how it should be presented.
[0081] Example 7 Based on the above embodiments, in this embodiment, the specific steps for obtaining the target intent template include: Based on the colloquial trigger word library, the intent data, and the presentation data, a candidate entity list is generated; for example, Team 1 will be mapped to the comprehensive mining team 1 and tunneling team 1 in the knowledge graph.
[0082] The candidate entity list and the composite intent knowledge graph are matched to obtain the candidate entity graph; the local graphs corresponding to the candidate entity list are obtained respectively, such as: Comprehensive Mining Team 1: Equipment A, Equipment B, Person in Charge Zhang San, Project P1; Tunneling Team 1: Equipment C, Equipment S, Person in Charge Li Si, Project P4. By utilizing the cosine similarity of the embedding vectors, the semantic similarity between the contextual semantic representation and the candidate entity graph (entity attributes, associated entities) is obtained to acquire the target data. If the data also includes coal mining data, the graph pattern matching score of the fully mechanized mining team will be higher, thereby achieving disambiguation. Finally, the mentions will be linked to globally unique entities.
[0083] The target data is mapped to the joint intent understanding model to obtain the target intent template.
[0084] Example 8 Based on the above embodiments, in this embodiment, if there are information gaps in the target intent template, the step of obtaining the scheduling instruction set based on the target intent template further includes: Construct a reasoning rule base, which includes reasoning rules for different reasoning types, including data type-control adaptation reasoning, entity type-default area reasoning, spatial semantics-associative presentation reasoning, role / position-view preference reasoning, and status-emergency presentation reasoning. Data type - control adaptation inference: Applicable control missing, inference path: query attribute (e.g., concentration) - query attribute has data characteristics - data type (e.g., time-varying numerical type) - best visualization method of data type - control type (e.g., trend curve control); Entity Type - Default Region Inference: Applicable region missing, inference path: Query entity (e.g., person) - Query the type to which the entity belongs - Entity type (e.g., person) - Default display area of entity type - Screen area; Spatial semantics - association presentation reasoning: applicable to missing associations, reasoning path: query entity (e.g., sensor A) - query entity is located in - region X (e.g., return airway of the second face of the fully mechanized mining) - monitoring method of region X - camera list (e.g., camera 01, camera 02). Role / Position - View Preference Inference: Applicable area missing; Inference path: Operator role (e.g., security inspector) - Operator role's preferred layout - Layout template (e.g., area monitoring priority layout). Status-Emergency Presentation Reasoning: Applicable to missing associations. Reasoning path: Query entity (e.g., sensor A) - Query entity has status - Status value (over-limit alarm has been triggered) - Triggering presentation method of status value - Emergency layout (e.g., central main screen preemption + alarm association pop-up).
[0085] Based on the inference rule base, the target intent template, and the target data, the missing information is obtained; Based on the missing information and the target intent template, the scheduling instruction set is obtained.
[0086] For example, a user command might be: "Help me see where the team members are right now, and also take a look at the situation on their work surfaces."
[0087] Parse the intent data: idea Figure 1 : Query the {current location} of {a team of people}. The displayed control can be inferred to be a GIS map, but the region is unknown (region is missing).
[0088] Secondary Intent 2: Query the {monitoring video} of {a team's work surface}. The displayed control is clearly a video, but the area is unclear (area missing), and it is inconsistent with the main idea. Figure 1 There is a strong correlation.
[0089] Inference is performed based on an inference rule base. For the idea Figure 1 reasoning: Starting from Entity 1, confirm its entity type by querying the type to which it belongs: Organizational Unit; Execution Entity Type - Default Region Reasoning: The default presentation region of an organizational unit is associated with the central main screen area (because personnel distribution is a core concern for scheduling). Conclusion 1: The idea Figure 1 The target area is filled in as the central main screen area.
[0090] Secondary Intent 2 Reasoning: Starting from Entity Team 1, the system navigates to the second face of the comprehensive mining area by querying the entity's location in the attribute field. Then, through the monitored relationship, it finds the associated camera 01 and camera 02.
[0091] Conclusion 2: The query entity of secondary intent 2 is automatically completed to these two specific cameras.
[0092] Step 3: Association and Spatial Layout Reasoning The idea Figure 1 Treating the auxiliary intention 2 as a complex scenario, we will perform overall layout reasoning: idea Figure 1 (Personnel location GIS map) is bound to the central main screen area.
[0093] Secondary Intent 2 (Workface Monitoring Video) due to the main idea Figure 1 The core area of the fully mechanized mining area has a strong semantic correlation with the monitored area, triggering spatial semantic-association presentation reasoning.
[0094] Bind these two surveillance videos to the auxiliary screen area on the right, and establish a business collaboration display relationship with the central main screen area.
[0095] Final complete presentation strategy: The central main screen area features a GIS map control, with the data source being the real-time location of a team of personnel.
[0096] Right auxiliary screen area: Dual-screen video control, with the data source being the RTSP streams from camera 01 and camera 02.
[0097] Association: The central area and the right side area are bound together as a team - location - monitoring information group.
[0098] Furthermore, after the reasoning is completed, the generated voice broadcast is no longer a simple repetition of the instructions, but includes a confirmation of the reasoning result: the location of a team of personnel has been displayed on the main screen, and the monitoring screen of their work area has been brought up on the right side of the screen.
[0099] Example 9 Based on the above embodiments, this embodiment also provides a voice dispatch and control system for multi-service linkage in mines, the system comprising: Graph Unit: Used to construct a composite intent knowledge graph in the mining field, and to obtain a joint intent understanding model based on the composite intent knowledge graph; Start-up unit: Based on the pre-trained mine background noise suppression model, it acquires a specific wake-up word, switches the voice acquisition device from sleep mode to working mode, acquires the voiceprint features of the current operator, and starts directional sound reception; Instruction unit: used to convert the collected speech into a text sequence, input the text sequence into the joint intent understanding model, and obtain a scheduling instruction set; Execution unit: used to execute the scheduling instruction set, acquire multi-source heterogeneous service data, and drive the dynamic reorganization of the physical large screen screen; Feedback unit: used to render the multi-source heterogeneous business data to the physical screen and generate voice broadcast feedback simultaneously.
[0100] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0101] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A voice dispatching and control method for multi-service linkage in mines, characterized in that, The method includes: S1. Construct a composite intent knowledge graph in the mining field, and obtain a joint intent understanding model based on the composite intent knowledge graph; S2. Based on the pre-trained mine background noise suppression model, a specific wake-up word is obtained to switch the voice acquisition device from sleep mode to working mode, and the voiceprint features of the current operator are obtained to enable directional sound reception; S3. Convert the collected speech into a text sequence, and input the text sequence into the joint intent understanding model to obtain a scheduling instruction set; S4. Execute the scheduling instruction set, acquire multi-source heterogeneous service data, and drive the physical large screen to dynamically reassemble the screen. S5. Render the multi-source heterogeneous business data onto the physical screen and generate voice broadcast feedback simultaneously.
2. The voice dispatching and control method for multi-service linkage in mines according to claim 1, characterized in that, The specific steps of S1 include: S101: Obtain all mining business entities and establish a colloquial thesaurus for each of the mining business entities; Construct queryable attributes for each of the mine-wide business entities, map the queryable attributes to the query interface of the backend business system, and construct triples based on the mine-wide business entities, the queryable attributes, and the query interface; Construct conditional slots for the queryable attributes, wherein the conditional slots include slot type, colloquial trigger word library and normalization transformation rules, and the slot type includes time range, spatial region, entity status, threshold comparison and identity identifier; Based on the intent of generating conditional combinations for the aforementioned conditional slots; Based on the combined conditions, filter parameters for the query interface are generated. Based on the triples, a semantic association layer of entity-attribute-query interface is constructed. S102: The physical screen is virtualized into a computable canvas, and the logical areas of the canvas and the area attributes of the logical areas are constructed. A preset visual control is defined, and the control properties of the visual control are constructed. The control properties include the data source types that can be bound and the minimum set of parameters required. Based on the logical region, the visual control, and the control-region adaptation rules, a dynamic screen layout template is constructed. Based on the aforementioned dynamic screen layout template, a knowledge layer for presenting large screen logical areas and visual controls is constructed. S103: Construct composite intent templates for different mine scheduling scenarios, each composite intent template containing a query intent and a presentation intent; Based on the semantic association layer, the presentation knowledge layer, and the composite intent template, the composite intent knowledge graph is constructed.
3. The voice dispatching and control method for multi-service linkage in mines according to claim 2, characterized in that, Before performing S4, the method further includes: Construct an occupancy status table for the logical area, which includes the currently occupied display task, the current priority level of the task, the timestamp when the task started occupying the area, whether the area can be preempted, and the core business entity associated with the currently displayed content. Obtain the logical relationships between the logical regions to obtain a cross-regional relationship diagram; Build a global history stack; Based on the occupancy status table, the cross-regional relationship diagram, and the global display history stack, a multi-dimensional digital twin model of the physical large screen display status is constructed. The scheduling instruction set is input into the multidimensional digital twin model to obtain the conflict results; The conflict results are matched with the conflict resolution rule base to obtain the target resolution rule, and an optimized scheduling instruction set is obtained based on the target resolution rule. The conflict resolution rule base includes resolution strategy rules under different conflict scenarios.
4. The voice dispatching and control method for multi-service linkage in mines according to claim 3, characterized in that, The resolution strategy rules include: Rule 1, absolute priority preemption: emergency preemption, alarm-type tasks have the highest priority, directly preempting the target area of non-alarm-type tasks; Rule 2, Content Collaboration and Coexistence: In-situ attribute switching, without creating a new window, directly switches the data source attributes of the controls bound to the current area, or overlays a temporary floating attribute panel on top of it; Rule 3, Intelligent Space Allocation: Flexible degradation allocation, automatically assigning new screens to the available candidate area with the highest matching value with the content type, and adaptively adjusting the window size to the optimal; Rule 4, Information Density Overload: Temporary Overlay and Fade-out. New content is overlaid on the existing content with the highest relevance in the form of a semi-transparent floating window, and automatically fades out after a preset time without operation. Rule 5, Context-Aware Clearing: Associated cascading clearing and focusing. When clearing an area, it automatically identifies and clears auxiliary screen areas that are business-related to that area.
5. A voice dispatching and control method for multi-service linkage in mines according to claim 4, characterized in that, The specific steps for matching the conflict results with the conflict resolution rule base to obtain the target resolution rule include: Based on the conflict results, a conflict feature vector is constructed, which includes request features, current occupancy features, and global situation features. Each of the aforementioned resolution strategy rules is converted into a conditional reasoning decision tree, and each rule node of the conditional reasoning decision tree contains a conditional expression and an action instruction; Based on the priority level, the conflict feature vector and the conditional reasoning decision tree are matched to obtain the target resolution rule.
6. A voice dispatching and control method for multi-service linkage in mines according to claim 5, characterized in that, If there are two or more resolution strategy rules with the same priority level and whose conditional expressions all satisfy the conditions, then the method further includes: Based on the resolution strategy rules, candidate rules are obtained, action instructions for the candidate rules are acquired, and a candidate operation list is obtained. A real-time state snapshot is obtained based on the aforementioned multidimensional digital twin model; Input the candidate operation list into the screen digital twin sandbox to obtain a snapshot of the predicted state; The operation cost is obtained by comparing the real-time state snapshot with the predicted state snapshot. Obtain the resolution strategy rule corresponding to the minimum value of the operation cost, and obtain the target resolution rule; The operational costs include visual stability costs, cognitive continuity costs, and information relevance costs. The visual stability cost includes window creation cost, window closing cost, and window movement / scaling cost; The cognitive continuity costs include entity-control unbinding costs, attribute-control in-situ switching costs, and data flow interruption and continuation costs. The cost of information correlation includes the cost of creating information silos and the cost of destroying information clusters.
7. A voice dispatching and control method for multi-service linkage in mines according to claim 6, characterized in that, The specific steps of S3 include: The text sequence is input into a joint intent network, which includes a shared encoder and parallel business operation intent decoders and screen presentation intent decoders, and the business operation intent decoders and screen presentation intent decoders are connected by cross attention. Based on the shared encoder, the text sequence is parsed to obtain several minimum input units and the contextual semantic representation of the minimum input units; Based on the business operation intent decoder and screen presentation intent decoder, the smallest input unit is parsed to obtain intent data and presentation data; The intent data and the presentation data are mapped to the joint intent understanding model to obtain the target intent template; the scheduling instruction set is obtained based on the target intent template.
8. A voice dispatching and control method for multi-service linkage in mines according to claim 7, characterized in that, The specific steps to obtain the target intent template include: Based on the colloquial trigger word library, the intent data, and the presentation data, a candidate entity list is generated; The candidate entity list and the composite intent knowledge graph are matched to obtain the candidate entity graph; Obtain the semantic similarity between the context semantic representation and the candidate entity graph to obtain the target data; The target data is mapped to the joint intent understanding model to obtain the target intent template.
9. A voice dispatching and control method for multi-service linkage in mines according to claim 8, characterized in that, If the target intent template has information gaps, the step of obtaining the scheduling instruction set based on the target intent template further includes: Construct a reasoning rule base, which includes reasoning rules for different reasoning types, including data type-control adaptation reasoning, entity type-default area reasoning, spatial semantics-associative presentation reasoning, role / position-view preference reasoning, and status-emergency presentation reasoning. Based on the inference rule base, the target intent template, and the target data, the missing information is obtained; Based on the missing information and the target intent template, the scheduling instruction set is obtained.
10. A voice dispatching and control system for multi-service linkage in mines, characterized in that, The system includes: Graph Unit: Used to construct a composite intent knowledge graph in the mining field, and to obtain a joint intent understanding model based on the composite intent knowledge graph; Start-up unit: Based on the pre-trained mine background noise suppression model, it acquires a specific wake-up word, switches the voice acquisition device from sleep mode to working mode, acquires the voiceprint features of the current operator, and starts directional sound reception; Instruction unit: used to convert the collected speech into a text sequence, input the text sequence into the joint intent understanding model, and obtain a scheduling instruction set; Execution unit: used to execute the scheduling instruction set, acquire multi-source heterogeneous service data, and drive the dynamic reorganization of the physical large screen screen; Feedback unit: used to render the multi-source heterogeneous business data to the physical screen and generate voice broadcast feedback simultaneously.