Emergency video monitoring scheduling method based on large model
Through large-scale model algorithms, video surveillance tag management and voice command recognition have been solved, and the number of video surveillance and complex management in the emergency command hall has been achieved, fast and accurate video resource retrieval and real-time data display have been achieved, and work efficiency has been improved.
Patent Information
- Application Number
- CN202510500220.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-01
AI Technical Summary
In the emergency command hall, the number of video surveillance is large, the management level is complex, and the naming is chaotic, which leads to high retrieval difficulties, slow real-time response, strong manual dependence, and difficulty in quickly and accurately understanding the leadership's intentions and retrieving the required resources.
The large-model algorithm is used to manage video surveillance tags, identify user voice commands, analyze intentions, realize intelligent scheduling, lower operation thresholds, provide voice control functions, and simplify operation processes.
It improves the response speed and work efficiency of video surveillance, reduces operation difficulty, enables more users to use the monitoring system efficiently, and realizes accurate scheduling of video resources and real-time data display.
Smart Images

Figure CN120238631A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video scheduling, and particularly to an emergency video monitoring and scheduling method based on a large model. Background Art
[0002] With the increasing demands for emergency safety and supervision, more and more functions in application systems, increasing data types and data volumes, problems such as difficult video monitoring and scheduling, difficult linkage of application systems, and difficult data information query have emerged. For the emergency command hall scenario, there are a large number of application systems, video monitoring, and business data that need to be operated and used. It is very difficult for business personnel to quickly and accurately understand the leadership's intentions and find the required resources from them. This has brought the following troubles to the staff and leaders in the emergency command hall.
[0003] Large number of videos: With the increasing demands for safety and supervision, the number of monitoring videos is constantly increasing;
[0004] High retrieval difficulty: There are many levels in video management, the monitoring naming is chaotic, and it is difficult to understand the user's query intentions;
[0005] Slow real-time response: In case of abnormal situations, it is difficult to make an immediate response, and there is a strong dependence on manual operation.
[0006] Therefore, for those skilled in the art, how to design a method that can quickly identify and retrieve videos according to user needs is an urgent problem to be solved. Summary of the Invention
[0007] By introducing a large model algorithm, the present invention can automatically analyze monitoring videos, quickly identify abnormal behaviors and events, reduce the workload of manual scheduling and monitoring, and improve the response speed; the voice control function of the present invention can simplify the operation process, enabling monitoring personnel to quickly schedule and view monitoring images while handling other tasks, improving the overall work efficiency; in addition, the present invention lowers the operation threshold: providing a simple and easy-to-use operation method for users who are not familiar with the technology, enabling more people to efficiently use the monitoring system.
[0008] To achieve the above object, the technical solution adopted by the present invention is: providing an emergency video monitoring and scheduling method based on a large model, including the following steps:
[0009] Step S101, perform label management on video monitoring and establish a point label information library;
[0010] Step S102, identify the user's voice command and perform information extraction;
[0011] Step S103, analyze the user's intention based on the intention information library;
[0012] Step S104, perform intelligent scheduling on video monitoring according to the user's intention.
[0013] Preferably, in step S101, establishing a point location tag information library includes the following steps:
[0014] Step S1011, point location identification definition: Provide the name and description of each point location, understand the location and use of the point location; Assign a unique identification code to each point location for quick identification of the point location;
[0015] Step S1012, point location attribute sorting: Define the type of the point location, record the physical location of the point location and technical parameters related to the point location;
[0016] Step S1013, point location management: Update the attribute, type and location information of the point location;
[0017] Step S1014, tag creation: Define tags according to the characteristics of the monitored point locations and business requirements;
[0018] Step S1015, continuous maintenance and update: Continuously maintain and expand the point location tag library using new data sources and user feedback to adapt to the development of application needs.
[0019] Preferably, in step S102, based on the semantic understanding ability of the large model, summarize and analyze the content in the voice command in the video dispatching scenario of the command center, and accurately fill it according to the slot system of administrative region - venue name - venue ontology - venue point location - point location ontology - monitoring attribute.
[0020] More preferably, the specific steps of summarizing and analyzing the content in the voice command include front - end voice processing and back - end voice processing, and establishing a corpus information library;
[0021] Among them, the content of establishing the corpus information library includes:
[0022] Corpus information classification: Classify the vocabulary according to the characteristics of the corpus in the field of emergency management;
[0023] Initial word library establishment: Establish an initial corpus by collecting emergency rescue data resources;
[0024] NLP extraction: Use text mining technology to discover new corpus vocabulary from text data;
[0025] Annotation rule formulation: Formulate annotation rules, including part - of - speech annotation, syntactic analysis, and entity recognition work to annotate the corpus information;
[0026] Conduct multiple rounds of quality control inspections on the processed text;
[0027] Review and supplement: Manually review the collected corpus vocabulary, remove inappropriate or incorrect entries, and update and supplement at the same time.
[0028] Preferably, the front - end voice processing function includes endpoint detection and noise cancellation.
[0029] Preferably, the back - end voice processing function includes a vocabulary library, continuous speech recognition, intelligent punctuation addition, confidence output, multiple recognition results, multi - slot recognition, hot - word recognition, and recognition logs.
[0030] Preferably, the confidence reflects the credibility of the recognition result. By using recall rate \(P\) Recall , accuracy rate \(P\) Precision , harmonic mean \(P\) average and Hamming loss function \(Ham\) loss as evaluation indicators for judgment, specifically as follows:
[0031]
[0032] Among them, \(N\) tp represents the number of samples predicted as \(p\) and the true class is \(p\); \(N\) fp represents the number of samples predicted as \(p\) and the true class is not \(p\); \(N\) ffp represents the number of samples predicted not as \(p\) and the true class is not \(p\); \(N\) represents the total number of samples, \(K\) represents the total number of labels, \(y\) and respectively represent the true value and the predicted value corresponding to the \(i\) - th sample, and XOR is the exclusive - OR logical operation.
[0033] Preferably, in step S103, by establishing an intention recognition library, integrating and analyzing the user's intention, extracting key information from the speaker's expression, judging the speaker's needs, and converting them into instructions according to the needs.
[0034] Preferably, establishing the intention recognition library also includes establishing an intention recognition library model, using artificial intelligence algorithms to train a classification model to recognize the intention of the user's speech, and supporting in the video dispatching scenario of the command center, synthesizing the feedback content of the intelligent assistant into speech for playback, realizing two - way voice interaction between the user and the intelligent dispatching system.
[0035] Preferably, in step S104, according to the analyzed user's intention, commanding and dispatching the on - duty personnel to quickly locate the video source and understand the on - site situation in real - time.
[0036] Compared with the prior art, the present invention also has the following advantages:
[0037] 1. By obtaining and processing a large amount of video data in real - time, setting the AIAgent video dispatching assistant with strong data integration capabilities, realizing the convergence, processing, and display of real - time data, and providing timely information support for decision - makers.
[0038] 2. Precise Video Scheduling: In dealing with emergencies, video surveillance is a very important source of information. Through the method of the present invention, by identifying the needs of the staff in the command center, relevant video resources can be quickly retrieved to achieve real-time online viewing.
[0039] 3. Human-Computer Interaction Optimization: To improve the work efficiency of the staff and reduce the operation difficulty, a human-computer interaction function is also set up to simplify the operation process and provide various interaction methods such as voice and text.
[0040] In summary, the present invention includes aspects such as real-time information acquisition and processing, intelligent question answering and information retrieval, precise video scheduling, system call and collaborative linkage, and human-computer interaction optimization, providing comprehensive support for the intelligent upgrading of governments or enterprises. Description of the Drawings
[0041] Figure 1 It is a flowchart of a method for emergency video surveillance and scheduling based on a large model of the present invention.
[0042] Figure 2 It is an operation flowchart of a preferred embodiment of the present invention.
[0043] Figure 3 It is a flowchart for opening video surveillance of the present invention. Detailed Embodiment
[0044] Please refer to Figure 1 As shown, the present invention relates to a method for emergency video surveillance and scheduling based on a large model, including the following steps:
[0045] Step S101, perform label management on video surveillance and establish a point location label information library;
[0046] Label management is an important function in resource management. By adding specific labels to resources, the classification, retrieval, and allocation of resources become more efficient and accurate. Labels can be any words or phrases describing the characteristics of resources, such as "scenic spots", "administrative regions", etc. Labels can also help us better understand the usage of resources for more effective resource planning and optimization. The present invention supports functions such as adding, deleting, modifying, and querying labels, supports label classification, and classifies and manages labels to facilitate users to search for and use labels.
[0047] The point location label information library classifies, organizes, and stores the relevant information of the monitoring points. The construction content focuses on efficiently and accurately recording and utilizing the information of each point location to support system monitoring, data analysis, and decision-making.
[0048] Establishing the point location label information library includes the following steps:
[0049] Step S1011, Point Identification Definition: Provide the name and description of each point, understand the location and usage of the point; assign a unique identification code to each point for quick identification of the point;
[0050] Step S1012, Point Attribute Sorting: Define the type of the point, record the physical location of the point and the technical parameters related to the point;
[0051] Step S1013, Point Management: Update the attribute, type and location information of the point;
[0052] Step S1014, Label Creation: Define labels according to the characteristics of the monitored points and business requirements;
[0053] Step S1015, Continuous Maintenance and Update: Continuously maintain and expand the point label library using new data sources and user feedback to adapt to the development of application needs.
[0054] Step S102, Identify the user's voice command and perform information extraction;
[0055] Based on the semantic understanding ability of the large model, summarize and analyze the content in the voice command in the video dispatching scenario of the command center, and accurately fill it according to the slot system of administrative region - venue name - venue ontology - venue point - point ontology - monitoring attribute.
[0056] The voice recognition ability engine of the present invention: Supports real-time recognition of the user's voice command into text in the video dispatching scenario of the command center. The problem to be solved by voice recognition technology is to enable the machine to "understand" human speech and "extract" the text information contained in the speech, which is equivalent to installing "ears" on the machine to make it have the function of "being able to listen".
[0057] The specific steps for summarizing and analyzing the content in the voice command include front-end voice processing and back-end voice processing, and establishing a corpus information library;
[0058] The front-end voice processing function includes endpoint detection and noise cancellation.
[0059] Front-end voice processing refers to using signal processing methods to perform preprocessing such as detecting and denoising the speaker's voice to obtain the voice most suitable for the recognition engine to process. The main functions include:
[0060] a. Endpoint detection
[0061] Endpoint detection is the process of analyzing the input audio stream to determine the start and end of the user's speech. Once it is detected that the user starts speaking, the voice starts flowing to the recognition engine until it is detected that the user stops speaking. This way enables the recognition engine to start the recognition process while the user is speaking.
[0062] b. Noise cancellation
[0063] It has efficient noise cancellation capabilities to meet the requirements of users in a wide variety of application environments.
[0064] The back-end speech processing functions include a vocabulary library, continuous speech recognition, intelligent punctuation addition, confidence output, multiple recognition results, multi-slot recognition, hot word recognition, and recognition logs.
[0065] (2) Back-end recognition processing
[0066] The back-end recognition processing recognizes the speaker's speech to obtain the most suitable results. The main features are:
[0067] a. Robust recognition function with a large vocabulary and independent of the speaker
[0068] The system meets the recognition requirements of a large vocabulary and speaker independence, and can support a vocabulary size of tens of thousands of grammar scales;
[0069] And it can adapt to application environments of different ages, regions, populations, channels, terminals, and noise environments.
[0070] b. Continuous speech recognition
[0071] Continuous speech recognition means being able to convert any speech spoken by the user into corresponding text information, supporting common sentence dictation, and having a high recognition accuracy for commonly used conversations in daily use.
[0072] c. Intelligent punctuation addition
[0073] Continuous speech recognition supports intelligent prediction of Chinese punctuation. Using a super-large language model, it intelligently predicts the dialogue context of the recognized result sentence, providing intelligent sentence segmentation and prediction of punctuation marks.
[0074] d. Confidence output
[0075] The confidence reflects the credibility of the recognition result. The speech recognition engine can carry the confidence of the recognition result when returning the recognition result, and the application program can analyze and perform subsequent processing based on the value of the confidence.
[0076] e. Multiple recognition results
[0077] Also known as multi-candidate technology, in some recognition processes, the recognition engine can return multiple recognition results that meet the conditions to the application program through the results of confidence judgment, rather than a single result. The recognition system provides a list of possible recognition results and arranges them in descending order of confidence results.
[0078] f. Multi-slot recognition
[0079] The slot in speech recognition represents a keyword, that is, during a conversation, multiple keywords contained in the speaker's speech can be recognized, which can improve the efficiency of speech recognition applications and enhance the user experience.
[0080] g. Hotword recognition
[0081] Hotword recognition enables a speech recognition application to detect a specific word or phrase while the speaker is speaking. When the speaker says this phrase, the recognition engine will return control to the application. Using this function in the application allows the recognizer to listen to the input speech in the background until the user says a specific phrase to make a request and then interact with the user.
[0082] h. Recognition log
[0083] The log of speech recognition plays a very important role in the system. This log records the input audio, the loaded grammar, the intermediate results of the recognition process, the recognition process of the recognition module, various parameters used in the recognition, the recognition results, and the system environment information at that time.
[0084] The content of establishing the corpus information library includes:
[0085] Corpus information classification: Classify the vocabulary according to the characteristics of the corpus in the field of emergency management;
[0086] Initial word library establishment: Establish an initial corpus by collecting emergency rescue data resources;
[0087] NLP extraction: Use text mining technology to discover new corpus vocabulary from text data;
[0088] Annotation rule formulation: Formulate annotation rules, including part-of-speech annotation, syntactic analysis, and entity recognition to annotate the corpus information;
[0089] Conduct multiple rounds of quality control checks on the processed text;
[0090] Review and supplementation: Manually review the collected corpus vocabulary, remove inappropriate or incorrect entries, and update and supplement them at the same time.
[0091] In order to improve the accuracy of speech recognition, the present invention also establishes a corpus information library. Customize personalized vocabulary and phrases according to the industry field and store them in the corpus information library, mainly including natural disaster vocabulary such as earthquakes, tsunamis, floods, typhoons and hurricanes, droughts, etc.; man-made disaster vocabulary such as explosions, fires, traffic accidents, chemical leaks, etc.; emergency response action vocabulary such as evacuation, search and rescue, post-disaster reconstruction, supply and replenishment, shelter and refuge, etc.; rescue supplies and equipment, shelters, life jackets, food and drinking water, tents, etc.
[0092] Among them, the confidence reflects the credibility of the recognition result. By adopting the recall rate P Recall , the accuracy rate P Precision , the harmonic mean P average and the Hamming loss function Ham loss as the evaluation indicators for judgment, the specific details are as follows:
[0093]
[0094] Among them, N tp represents the number of samples predicted as p and the true category is p; N fp represents the number of samples predicted as p and the true category is not p; N ffp represents the number of samples predicted not as p and the true category is not p; N represents the total number of samples, K represents the total number of labels, y and respectively represent the true value and the predicted value corresponding to the i-th sample, and XOR is the exclusive OR logical operation. Through this calculation method, the semantic understanding ability of the large model for the user's speech statement is continuously trained, updated, and iterated until the optimal large model is output, improving the recognition accuracy.
[0095] In terms of the P Recall indicator, the P Recall identified by the method of the present invention can reach a relatively high level of more than 70%. Due to the trade-off between precision and recall, P Precision although it does not reach a relatively high level, but P average also achieves the optimal performance. Compared with other existing technologies, the method of the present invention reduces the Hamming loss by 1 to 3 percentage points and obtains a lower Hamming loss value. This indicates that in reducing the error prediction rate, it proves that the method of the present invention has a significant advantage over the existing technologies in improving the overall performance of multi-label classification.
[0096] The method of the present invention not only considers the semantic relationship between sample labels, but also endows the labels with richer and more accurate semantic representations through the semantic meaning of the labels by the large model. At the same time, the method of this article also fully integrates the artificial intelligence to recognize the intention of the user's speech, thereby on the other hand, feeding back the training effect of semantic recognition and improving the accuracy of semantic recognition and intention recognition.
[0097] Step S103, analyze the user's intention based on the intention information library;
[0098] By establishing a consciousness graph recognition library, integrating and analyzing the user's intention, extracting key information from the speaker's statement, judging the speaker's needs, and converting them into instructions according to the needs.
[0099] Building an intent recognition library also includes building an intent recognition library model, training a classification model using artificial intelligence algorithms to recognize the intent of the user's speech, and supporting the synthesis of the feedback content of the intelligent assistant into speech for playback in the video dispatching scenario of the command center, realizing two-way voice interaction between the user and the intelligent dispatching system.
[0100] The content of the intent library includes information such as intent name, intent description, intent reply, feedback operation, etc. The construction content includes:
[0101] 1. Defining intents and sample collection: Analyze the possible needs of users, and based on this, define the intents that users may express; collect various expressions of the same intent by different users.
[0102] 2. Annotation and classification: Manually annotate the intent samples to ensure that each sample is correctly classified into the corresponding intent category.
[0103] 3. Intent pattern creation: According to the labeled data, create or extract patterns that can summarize each intent, and these patterns include keywords, phrase structures, syntactic features, etc.
[0104] 4. Building an intent recognition model: Use artificial intelligence algorithms to train a classification model to recognize the intent in the user input.
[0105] 5. Continuous maintenance and update: Continuously maintain and expand the intent knowledge base using new data sources and user feedback.
[0106] The intent recognition algorithm of the present invention supports the recognition of the user's instruction intent, instruction parameter extraction, and label matching in the video dispatching scenario of the command center, so as to support the system to match the conversation to reply to the user. The intent management implements a two-level management system, and each category has different types of sub-intents for configuration. The intent flow in all instructions is controlled by the intent recognition and multi-round dialogue algorithms. The main functions are as follows:
[0107] 1. Localized deployment.
[0108] 2. Intent matching algorithm: Qualitatively analyze the voice instructions, and map the instructions to the predefined intent categories for correct response, realizing accurate matching of the user's intent. Among them, the ACC reaches 95.4%, the recall reaches 96.4%, the precision reaches 95.6%, the f1_score reaches 95.7%, and the inference duration does not exceed 3 seconds.
[0109] 3. Parameter extraction algorithm: A technical process of extracting the most representative and key features from the voice instructions to support subsequent analysis, processing, or decision-making. Among them, the algorithm is evaluated on the validation dataset, the recall reaches 96.19%, the precision reaches 97.2%, and the f1_score reaches 96.69%.
[0110] 4. Label matching algorithm: A technology that corresponds the input voice command to predefined classification labels to achieve the purpose of identifying, classifying, or organizing data and realizing the goal of precise matching.
[0111] 5. Monitor and schedule sub-intentions. Define sub-intentions by venue type and construct them in combination with business scenario information. Sub-intentions are responsible for managing default points and customizing reply scripts.
[0112] After extracting and recognizing the user's voice command in step S102, match the data in the intention information library. After analyzing the user's intention through artificial intelligence AI in the intention information library, input the command into the interaction system and dispatch the video through artificial intelligence AI.
[0113] The present invention can not only communicate with users in text but also conduct voice interaction. The voice synthesis ability engine of the present invention: supports synthesizing the feedback content of the intelligent assistant into voice for playback in the video dispatch scenario of the command center. The voice synthesis ability mainly broadcasts the reply content given by the intelligent dispatch system to the user in voice, realizing two-way voice interaction between the user and the intelligent dispatch system.
[0114] The voice synthesis system integrated in the voice platform is a leading text-to-speech engine in the industry. It adopts advanced Chinese text, prosody analysis algorithms, and synthesis methods of large corpora, and the synthesized voice has approached the natural effect of real people. Its main functions are:
[0115] (1) High-quality voice, converting the input text into smooth, clear, natural, and expressive voice data in real time;
[0116] (2) High-precision text analysis technology, ensuring intelligent analysis and processing of out-of-vocabulary words (such as place names), polyphonic characters, special symbols (such as punctuation marks, numbers), prosodic phrases, etc. in the text;
[0117] (3) Support for multiple character sets, supporting the input of various character sets such as GB2312, GBK, Big5, Unicode, and UTF-8, as well as text information in various formats such as ordinary text and text with CSSML annotations;
[0118] (4) Multiple data output formats, supporting the output of voice data in various formats such as linear Wav, A / U rate Wav, and Vox with various sampling rates;
[0119] (5) Provide a pre-recorded synthesis template, using the pre-recorded voice of the speaker for the text in the synthesized text that conforms to the fixed components of the voice template, and using synthesized voice for non-fixed components;
[0120] (6) Voice adjustment function. The development interface provides various dynamic adjustment functions for synthesis parameters such as volume, speech rate, and pitch (fundamental frequency).
[0121] (7) Configuration and management tools. The synthesis engine provides tools for unified configuration and management, completing functions such as global parameter configuration, user dictionary, user rules, and customized resource package management.
[0122] (8) Effect optimization. The synthesis engine provides various methods for optimizing synthesis effects for actual application environments, represented by customized resource packages and CSSML.
[0123] Speech synthesis mainly provides the ability for intelligent assistants to perform machine automatic announcements. For the interactive content replied by the intelligent assistant to the user, speech synthesis will be used to enable the system to automatically announce it, optimizing the human-computer interaction experience.
[0124] Step S104, perform intelligent scheduling on video surveillance according to the user's intention.
[0125] The real-time monitoring and precise scheduling of surveillance videos for traffic, forest fire prevention, eagle eyes, etc. in the present invention can help the command and dispatch duty personnel quickly locate the key video sources, understand the on-site situation in real time, and take response measures in a timely manner, improving the emergency response ability and work efficiency.
[0126] The present invention can realize monitoring execution management and support zooming in / out of specific surveillance videos through voice commands. For example: "Zoom in on the third surveillance for me", "Zoom out on the first video for me", etc.; Here, the entire process will be introduced taking "Help me open the video surveillance by the sea" as an example, as Figure 2 shown. The large model of the present invention recognizes that "by the sea" belongs to "scene familiarity" and further obtains "video attribute" information, including: obtaining video data through the keyword of "affiliated institution"; obtaining video data through the keyword of "video name"; obtaining video data through the keyword of "scene label" (including label description); obtaining video data through the keyword of "address"; and then opening the video surveillance by merging the video data. The flowchart for opening the video is as Figure 3 shown.
[0127] The present invention can assist in monitoring and analyzing the online status of monitoring devices, enabling users to clearly understand the working status of each monitoring device, promptly detect and resolve issues such as device offline or malfunction, and ensure the stable operation of the monitoring system. It supports docking with public security, traffic, forest fire prevention, and eagle-eye monitoring videos to achieve real-time monitoring and precise scheduling, and helps the command and dispatch duty personnel quickly locate key video sources. It supports pushing the finally dispatched monitoring resources to the large screen, corresponding to the video monitoring function module. It supports recording the problems of each video dispatch, the dispatched monitoring resources, and the actual connectivity rate of the final monitoring resources to meet the self-learning and adaptive real-scene of the algorithm. It supports calculating the actual connectivity rate of the monitoring resources dispatched each time. The calculation formula is: connectivity success rate = actual connected quantity / dispatched monitoring resource quantity. It supports analyzing the monitoring dispatch logs using a large model, which can summarize the monitoring resources frequently dispatched through the Xiaoying Assistant within a specified time and generalize the monitoring resources that are often unable to be connected.
[0128] The above embodiments are only descriptions of the preferred embodiments of the present invention and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.
Claims
1. A large model-based emergency video surveillance scheduling method, characterized in that: The following steps are involved: Step S101, perform label management on video surveillance and establish a point label information library; Step S102, recognizing the user's voice command and extracting information; Step S103, analyzing the user's intention based on the intention information library; Step S104: Intelligently schedule video surveillance according to user intention.
2. According to the large model-based emergency video monitoring scheduling method of claim 1, it is characterized in that: In step S101, establishing a point tag information library includes the following steps: Step S1011, point identification definition: provide a name and description for each point to understand the location and purpose of the point; assign a unique identification code to each point to quickly identify the point; Step S1012, sorting out point attributes: defining the type of the point, recording the physical location of the point and technical parameters related to the point; Step S1013, point management: update the attributes, type and location information of the point; Step S1014, label creation: define labels according to the characteristics of the monitoring points and business requirements; Step S1015, continuous maintenance and updating: using new data sources and user feedback to continuously maintain and expand the point tag library to adapt to the development of application needs.
3. The large model-based emergency video surveillance scheduling method according to claim 1 is characterized in that: In step S102, based on the semantic understanding ability of the large model, the content in the voice command is summarized and analyzed in the command center video scheduling scenario, and accurately filled in according to the slot system of administrative area-venue name-venue body-venue point-point body-monitoring attribute.
4. The large model-based emergency video surveillance scheduling method according to claim 3 is characterized in that: The specific steps of summarizing and analyzing the content in the voice command include front-end voice processing and back-end voice processing, and establishing a corpus information database; The content of establishing the corpus information database includes: Corpus information classification: Classify vocabulary according to the characteristics of the corpus in the field of emergency management; Initial lexicon establishment: Establish an initial corpus by collecting emergency rescue data resources; NLP extraction: using text mining technology to discover new corpus vocabulary from text data; Annotation rule formulation: formulate annotation rules, including part-of-speech tagging, syntactic analysis, and entity recognition to annotate corpus information; Conduct multiple rounds of quality control checks on the processed text; Review and supplement: Manually review the collected corpus vocabulary, remove inappropriate or erroneous entries, and update and supplement them.
5. The large model-based emergency video surveillance scheduling method according to claim 4 is characterized in that: The front-end voice processing functions include endpoint detection and noise cancellation.
6. The large model-based emergency video surveillance dispatching method according to claim 4 is characterized in that: The backend voice processing functions include vocabulary library, continuous speech recognition, intelligent punctuation addition, confidence output, multiple recognition results, multi-slot recognition, hot word recognition and recognition log.
7. The large model-based emergency video surveillance dispatching method according to claim 6 is characterized in that: The confidence reflects the credibility of the recognition result, and is calculated by using the recall rate P Recall , accuracy P Precision , harmonic mean P average And the Hamming loss function loss The evaluation indicators used for judgment are as follows: Among them, N tp It is expressed as the number of samples predicted as p and the true category is p; N fp It is expressed as the number of samples predicted to be p and whose true category is not p; N ffp It is represented by the number of samples whose prediction is not p and whose true category is not p; N represents the total number of samples, K represents the total number of labels, y and They are respectively represented as the true value and predicted value corresponding to the i-th sample, and XOR is an exclusive OR logic operation.
8. The large model-based emergency video surveillance dispatching method according to claim 1 is characterized in that: In step S103, by establishing a consciousness map recognition library, integrating and analyzing the user's intentions, extracting key information from the speaker's expression, judging the speaker's needs, and converting them into instructions according to the needs.
9. The large model-based emergency video surveillance scheduling method according to claim 8, characterized in that: The establishment of the intent recognition library also includes establishing an intent recognition library model, using an artificial intelligence algorithm to train a classification model to recognize the intention of the user's speech, and supporting the synthesis of the intelligent assistant's feedback content into voice for playback in the command center video scheduling scenario, thereby realizing two-way voice interaction between the user and the intelligent scheduling system.
10. The large model-based emergency video surveillance dispatching method according to claim 1, characterized in that: In step S104, based on the analyzed intention of the user, the on-duty personnel are directed and dispatched to quickly locate the video source and understand the on-site situation in real time.
Citation Information
Patent Citations
Electric power information query and regulation function calling method based on voice and intention recognition
CN113314106A
Monitoring scheduling method, system and device and storage medium
CN113407771A
Quick video retrieval method, device and equipment based on natural language
CN117688207A
Scheduling instruction generation method and device, medium and program product
CN118824250A
Intelligent control method, system and equipment of railway comprehensive video monitoring system and medium
CN119052424A
Cited By
Downhole operation safety detection method, device and equipment
CN122200485A
Intelligent video monitoring scheduling system and method
CN122309804A