A method and system for multi-modal real-time interactive decision-making

Through the multimodal real-time interactive decision-making method, video and voice data are integrated, and the problems of untimely and inefficient information acquisition of traditional military decision-making systems are solved, real-time decision-making support and information integration are achieved, and military combat effectiveness is improved.

CN118262114BActive Publication Date: 2025-06-10XINGZHI INTELLIGENT (BEIJING) TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410384207.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-01
Publication Date
2025-06-10
Estimated Expiration
2044-04-01

AI Technical Summary

Technical Problem

Traditional military decision-making systems have problems such as fragmentation of information, poor real-time performance, high-intensity work and insufficient information richness, resulting in untimely and inefficient information acquisition, which cannot meet the ever-changing needs of the military field.

Method used

The multimodal real-time interactive decision-making method is adopted to obtain video data and speech signals, pre-processing, target recognition, semantic segmentation, visual significance detection and optical flow estimation, integrate video mining results and speech features, generate answer text and convert it into speech signals, and provide real-time decision support.

Benefits of technology

Real-time decision-making support is achieved, multi-source information is effectively integrated, the accuracy and adaptability of voice processing is improved, the accuracy and stability of video mining is enhanced, the work burden of commanders is reduced, and military combat effectiveness is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118262114B_ABST
    Figure CN118262114B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for multi-modal real-time interactive decision-making, belonging to the field of Internet technology. The method includes: acquiring video data and voice signals, preprocessing the video data to obtain preprocessed video data; performing object recognition on the preprocessed video data to obtain target objects; using a semantic segmentation network model to perform semantic segmentation on the preprocessed video data to obtain segmentation results; extracting important regions from the preprocessed video data through visual saliency detection technology; analyzing the important regions through an optical flow estimation method to obtain abnormal motion behaviors; fusing the target objects, segmentation results, and abnormal motion behaviors to obtain video mining results. The present invention also discloses a system for multi-modal real-time interactive decision-making. The present invention provides real-time decision support, and based on the multi-modal data supplemented and deduced by the large model, enables the commander to more quickly acquire, analyze, and understand key information, strengthening the real-time perception of the battlefield.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technologies, and particularly to a method and system for multi-modal real-time interactive decision-making. Background Art

[0002] With the increasing complexity and variability of the modern military operation environment, military decision-making requires more timely and comprehensive information support. Traditional military decision-making systems mainly rely on limited sensor data and manual intelligence analysis, suffering from problems such as untimely information acquisition and low efficiency. Therefore, there is an urgent need for a method and system for multi-modal real-time interactive decision-making to integrate voice and video technologies, provide more comprehensive and immediate intelligence support, and enhance the intelligence and automation of command decision-making.

[0003] Currently, in the interaction mode of traditional military assistance systems, the following problems mainly exist:

[0004] (1) Information fragmentation: The information obtained by sensors is in a fragmented state, which needs to be comprehensively analyzed and integrated. It requires manual access to various types of information for integration. Information such as text, video, and voice is independent of each other and difficult to fuse.

[0005] (2) Poor real-time performance: In the traditional mode, information is generally obtained manually or semi-automatically and then decisions are made, which takes a lot of time, affects the timeliness of decision-making, results in a lag in decision response, and cannot meet the ever-changing needs of the military field.

[0006] (3) High-intensity work: Commanders need to process multi-source information simultaneously, and even need to actively type and search, and consult various types of video, voice, and text content, and then give conclusions and decisions after thinking, which is very prone to information overload and fatigue.

[0007] (4) Insufficient information richness: For information monitored or generated from multiple sources, the mode is generally relatively single. During the analysis process, only the actually obtained information can generally be used, and it is impossible to generate or supplement relevant information to infer a more complete result, and the information is generally returned to the commander in pure text form, with poor intuitiveness.

[0008] In a fast-paced military scenario, commanders need to quickly acquire, analyze, and convey key information, including voice information from troops and intelligence personnel. Existing technologies are basically multiple independent modules, using different technologies or even different systems to separately acquire, store, and process different types of data such as text, voice, and images, and the efficiency of human-computer interaction is low, failing to exert the ability of multi-modal information processing, resulting in scattered information and being unable to quickly obtain the key points to make effective decisions. Since the military field requires real-time monitoring of videos to identify potential threats, if information processing is not timely, it may seriously affect the success rate of relevant operations.

[0009] General voice processing capabilities have the following main defects:

[0010] Lack of real-time performance: Existing speech segmentation and recognition technologies lack real-time performance in high-noise and dynamic environments, affecting the commander's immediate understanding of the battle situation.

[0011] Insufficient recognition capabilities for professional scenarios: Some existing technologies have low recognition accuracy for special accents or specific industry terms, and cannot be directly applied to voice input and generation in scenarios with a large number of industry terms such as military, resulting in a high misrecognition rate.

[0012] Difficulty in information integration: Existing methods fail to effectively extract key information from different data and integrate information from different voice sources, causing commanders to divide their attention to integrate multi-source information.

[0013] Unable to handle the problem of missing information: Existing methods can generally only recognize based on the currently acquired speech. If the noise is too loud or the signal is poor in the middle, the missing or poor-quality information of the speech generally cannot be effectively recognized, resulting in incoherent and incomplete speech.

[0014] General video processing capabilities have the following main defects:

[0015] Unstable target recognition: General recognition technology is sensitive to environmental changes such as lighting and weather, which limits the stable recognition of targets. However, in the military field, monitoring of targets requires that no suspicious targets be missed, and the error tolerance rate is very low. In addition, existing video detection technology is prone to distortion when processing targets with high-speed movement and complex movements, which limits the timely detection of abnormal behavior.

[0016] Inaccurate pixel semantic segmentation: Targets in military scenarios generally have anti-reconnaissance capabilities and are relatively concealed. Commonly used pixel semantic segmentation techniques often cannot accurately divide targets. It is necessary to enhance the information and learn the characteristics of targets in various states in order to achieve accurate identification.

[0017] The parts not captured are difficult to use: Existing image processing models can generally only extract and analyze based on the images that have been monitored. However, the images in military scenarios are more complex and changeable, and the targets move quickly, making it often difficult to capture the full picture of things. Traditional models cannot intelligently infer more image content that should be outside the screen, and sometimes it is not possible to obtain enough effective information based solely on the content on the picture. Summary of the invention

[0018] The object of the present invention is to provide an efficient multi-modal real-time interactive decision-making method and system.

[0019] To solve the above technical problems, the present invention provides a method for multi-modal real-time interactive decision-making, including the following steps:

[0020] Obtain video data and voice signals;

[0021] Preprocess the video data to obtain preprocessed video data;

[0022] Perform object recognition on the preprocessed video data to obtain target objects;

[0023] Use a semantic segmentation network model to perform semantic segmentation on the preprocessed video data to obtain segmentation results;

[0024] Extract important regions from the preprocessed video data through visual saliency detection technology;

[0025] Analyze the important regions through an optical flow estimation method to obtain abnormal motion behaviors;

[0026] Fuse the target objects, segmentation results, and abnormal motion behaviors to obtain video mining results;

[0027] Preprocess the voice signal to obtain preprocessed voice signal;

[0028] Extract features from the preprocessed voice signal to obtain voice features;

[0029] Perform voiceprint recognition and speech recognition on the voice features to obtain question texts;

[0030] Generate answer texts according to the question texts;

[0031] Convert the answer texts into voice signals to obtain answer voice signals;

[0032] Send the video mining results and answer voice signals to the user.

[0033] Preferably, it further includes the following steps:

[0034] Supplement the target objects according to the segmentation results or abnormal motion behaviors.

[0035] Preferably, using a semantic segmentation network model to perform semantic segmentation on the preprocessed video data specifically includes the following steps:

[0036] Based on the position and category information of the target objects and relevant deep learning algorithms, perform semantic segmentation on the preprocessed video data.

[0037] Preferably, the important region is the position of a certain target object or the region extracted from the segmentation results for a certain semantic category.

[0038] Preferably, the important regions are analyzed by an optical flow estimation method to obtain abnormal motion behaviors, which specifically include the following steps:

[0039] According to the target object or the segmentation result, analyze the motion of the objects in the important regions to obtain abnormal motion behaviors.

[0040] Preferably, the target recognition uses a region-based convolutional neural network or a single-stage detector;

[0041] The semantic segmentation network model uses a convolutional neural network or a fully convolutional neural network;

[0042] The preprocessing of the video data includes video frame extraction, inter-frame compensation processing, and video size unification processing.

[0043] Preferably, the preprocessed speech signal is subjected to feature extraction to obtain speech features, which specifically include the following steps:

[0044] The speech signal is transformed into a spectral representation in the time-frequency domain through short-time Fourier transform or its variant, and the Mel spectral coefficient features are extracted to obtain speech features.

[0045] Preferably, the speaker recognition uses a Gaussian mixture model or a deep learning model;

[0046] The speech recognition uses a hidden Markov model or a deep learning model;

[0047] The transformation of the speech signal uses a synthesis algorithm based on splicing units, a parameter generation model, or a deep learning model.

[0048] Preferably, the preprocessing of the speech signal includes background noise, volume equalization, removal of silent segments, and intelligent filling of missing segments.

[0049] The present invention also provides a multi-modal real-time interaction decision-making system, including:

[0050] An acquisition module for acquiring video data and speech signals;

[0051] A video preprocessing module for preprocessing the video data to obtain preprocessed video data;

[0052] A target recognition module for performing target recognition on the preprocessed video data to obtain target objects;

[0053] A semantic segmentation module for performing semantic segmentation on the preprocessed video data using a semantic segmentation network model to obtain a segmentation result;

[0054] An important region extraction module for extracting important regions from the preprocessed video data through visual saliency detection technology;

[0055] A video detection module, which is used to analyze important areas through an optical flow estimation method to obtain abnormal motion behaviors;

[0056] A fusion module, which is used to fuse the target object, the segmentation result and the abnormal motion behavior to obtain a video mining result;

[0057] A voice preprocessing module, which is used to preprocess voice signals to obtain preprocessed voice signals;

[0058] A feature extraction module, which is used to extract features from the preprocessed voice signals to obtain voice features;

[0059] A voiceprint and speech recognition module, which is used to perform voiceprint recognition and speech recognition on voice features to obtain question texts;

[0060] An answer generation module, which is used to generate answer texts according to the question texts;

[0061] A speech synthesis module, which is used to convert the answer texts into voice signals to obtain answer voice signals;

[0062] A sending module, which is used to send the video mining result and the answer voice signal to the user.

[0063] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0064] Real-time decision support: By integrating advanced speech processing and video mining technologies, the present invention provides real-time decision support. According to the multi-modal data supplemented and deduced by the large model, the commander can obtain, analyze and understand key information more quickly, strengthening the real-time perception of the battlefield.

[0065] Integration of multi-source information: The present invention can effectively integrate information from different voice and video sources, providing integrated multi-modal data to help the commander comprehensively grasp the battlefield situation and strengthen the comprehensive analysis of multi-source information.

[0066] Improve the accuracy and adaptability of speech processing: Through advanced speech segmentation, voiceprint recognition and speech recognition technologies, the present invention improves the accuracy of speech processing and at the same time provides stronger adaptability. It can process various accents, dialects and professional terms, and can generate the missing or low-quality parts of the speech through the large model, improving the quality and usability of voice information.

[0067] Enhance the accuracy and stability of video mining: The object recognition, pixel semantic segmentation and video detection technologies of the present invention improve the accuracy and stability of video mining. It can deduce all-round information from partial monitoring results, making it more adaptable to complex battlefield environments and providing more reliable object recognition and tracking capabilities.

[0068] Improving User Interaction Experience: Through advanced voice generation technology, the present invention improves the voice interaction experience and enhances the interaction efficiency and friendliness between the commander and the military intelligent system.

[0069] Self - adaptability and Intelligent Decision - making: The present invention has self - adaptability and can automatically adjust algorithm parameters according to different military scenarios and user requirements to achieve intelligent decision - making support, improving the flexibility and applicability of the system.

[0070] Reducing the Burden on the Commander: Through automation and intelligent technology, the present invention reduces the burden on the commander in information analysis and decision - making, enabling the commander to focus more on formulating strategies and guiding operations.

[0071] Enhancing Military Combat Effectiveness: Considering the above - mentioned beneficial effects, the present invention is expected to significantly enhance military combat effectiveness, enabling the commander to more advantageously cope with complex and changing battlefield environments and increasing the odds of victory on the battlefield. Brief Description of the Drawings

[0072] The following further details the specific implementation manners of the present invention in conjunction with the drawings.

[0073] Figure 1 It is a schematic diagram of the step - by - step process of video mining;

[0074] Figure 2 It is a schematic diagram of the collaborative work of three functional modules: target recognition, pixel semantic segmentation, and video detection;

[0075] Figure 3 It is a schematic diagram of the step - by - step process of voice processing;

[0076] Figure 4 It is a schematic diagram of the collaborative work of the voice signal processing sub - model with the voiceprint recognition and speech recognition sub - models respectively;

[0077] Figure 5 It is a schematic diagram of the process of a multi - modal real - time interaction decision - making method of the present invention. Specific Implementation Manner

[0078] Many specific details are set forth in the following description to facilitate a thorough understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific implementations disclosed below.

[0079] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0080] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0081] The following further describes the present invention in detail with reference to the accompanying drawings:

[0082] In view of the high requirements for real-time performance and reliability in the military field, the present invention proposes a method and system for multi-modal real-time interactive decision-making to make up for the deficiencies of the prior art in the real-time performance of information acquisition. By integrating advanced speech processing and video mining technologies, the present invention supports real-time interaction between users and intelligent military assistance systems, comprehensively captures key points in various types of information such as video, images, voices, and texts, and can, through multi-modal large model technologies, intelligently complement other information outside the currently acquired information range, so as to make more timely and effective judgments based on more comprehensive data, improve the intelligence level of the decision-making system, enable it to better adapt to the rapidly changing battlefield environment, effectively reduce the burden on commanders in information analysis and decision-making, and improve the command and decision-making efficiency.

[0083] The multi-modal information processing model mainly serves military scenarios and cooperates with the voice interaction system to assist commanders in making decisions and predictions quickly. The video mining model can effectively identify various types of video content in real-time shooting or historical archives, and intelligently segment, identify, and expand the images under each frame of the picture. Through the capabilities of the multi-modal large model trained with a large number of domain pictures, this method can effectively generate and expand the content of the current picture, extract the key information contained in the video data, automatically infer and generate more off-screen information, and introduce it into the knowledge base to fuse with other existing modal knowledge, comprehensively analyze the key information, and automatically generate multi-modal conclusions with pictures and texts. The commander can directly have a voice conversation with the business system and can support the completion of incomplete voices through the multi-modal large model to ensure the effectiveness of the instructions. When the system performs information retrieval, it will comprehensively consider multi-modal data and display the results in a multi-modal way to the user, providing the user with a more comprehensive information perspective and a more convenient interaction method, improving the efficiency and correctness of the commander's decision-making.

[0084] Specifically, multi-modal information mainly includes two major capabilities: video (image) processing and voice processing. The video mining model can detect and identify key objects, scenes, or events in the video through the video detection module, and can automatically generate and expand more abundant off-screen information through the multi-modal large model; it can extract key frames from the video sequence and perform semantic segmentation on each pixel in the video frame, classifying them into different semantic categories; it can intelligently analyze the video data, understand the scene content in the video, and achieve automatic recognition of the target of interest and automatic monitoring of abnormal behaviors; it can learn the visual features of the target object, perform accurate target detection and tracking in the video sequence, identify key information such as its category and attributes, and promptly feedback to the commander; it can, based on the large model learned from a large amount of historical data, intelligently analyze the information in the current picture and automatically complete the off-screen information not captured by the video. The voice processing model supports the commander to operate the business system by voice, and can identify the user's identity through voice characteristics for permission control and personalized services; the voice processing model can identify problems from the voice, extract key information, and convert it into corresponding data query instructions to achieve the exploration and query of multi-modal data; when the signal is poor or the noise is large, it can effectively complete the key information in the voice based on the multi-modal large model to ensure the integrity of the instructions; the voice processing model supports converting the text information generated by the system into a voice signal through voice conversion and returning it to the commander for quick interactive communication.

[0085] The functional requirements of the video mining model are described as follows: The functional requirements of the video mining model include video content analysis, feature extraction and representation learning, video detection and tracking, pixel-level semantic segmentation, and multimodal data fusion. It can infer more comprehensive content based on large models combined with existing information, realizing in-depth analysis and understanding of video data to provide rich multimodal information and accurate video content processing capabilities. The main functions involve three aspects: object recognition, pixel semantic segmentation, and video detection.

[0086] Object recognition supports efficiently and accurately locating and identifying predefined special objects from a large number of images, continuously detecting and tracking specific target objects. Pixel semantic segmentation can identify pixels of different categories in the target image and classify the feature pixel labels in the image. Video detection uses visual saliency detection technology to filter out relatively unimportant visual regions, thereby performing more targeted video anomaly detection on important regions to achieve more efficient detection and calculation. If the target in the original picture is not fully captured, this method can automatically complete the picture and identify the existing target or risk information points through the large model's understanding ability of a large number of military images.

[0087] The step process of the video mining model is as Figure 1 shown and includes the following steps:

[0088] (1) Data preparation and preprocessing: Collect and prepare the video dataset, including labeled objects, pixel semantic segmentation labels, and basic information of the video. Then, preprocess the video data, including video frame extraction, inter-frame compensation, video size unification, etc., for subsequent processing and analysis;

[0089] (2) Object recognition: Use multimodal generation technology to supplement and enhance the image information through a large model to improve the integrity of the picture, and use object detection algorithms, such as region-based convolutional neural network (R-CNN) or single-stage detector (SSD), to detect and locate objects in the video frames and further classify the detected objects in each video frame to identify the categories and attributes of the target objects;

[0090] (3) Pixel semantic segmentation: Use deep learning-related algorithms, including but not limited to convolutional neural network (CNN) or fully convolutional neural network (FCN), to perform semantic segmentation on the pixels in the video frames, assign them to different semantic categories, and post-process the segmentation results, including edge smoothing, pixel merging, etc., to obtain more accurate pixel-level semantic segmentation results;

[0091] (4) Video detection: Using visual saliency detection technology to identify important regions in the video, filtering out relatively unimportant visual regions to improve the efficiency and accuracy of video detection. Analyze the motion of objects in the video using optical flow estimation method, capture abnormal motion behaviors, and perform abnormal detection and calculation;

[0092] (5) Result fusion and output: Integrate the results of object recognition, pixel semantic segmentation, and video detection to obtain the complete video mining results, and store the mining results in the knowledge base to achieve the management and retrieval functions of multi-modal data. The specific objects recognized will be associated with the pixel blocks to which the objects belong. The abnormal motion behaviors detected in a specific frame of the video will also be associated with specific pixel blocks. If it is an abnormal motion of a specific object, it will be associated with the specific object.

[0093] The video mining model includes three functional modules: object recognition, pixel semantic segmentation, and video detection. These modules work together to identify, monitor, and extract effective information from the video, supporting the user's multi-modal information acquisition needs, as Figure 2 shown.

[0094] The combination of object recognition and pixel semantic segmentation can complement and work together to provide a more comprehensive and accurate understanding of video content: The object recognition model can provide the location and category information of the target object, helping the pixel semantic segmentation model to perform more accurate pixel-level classification within the target area, and focusing on generating reasoning for specific target areas to obtain more information; The pixel semantic segmentation model can provide richer semantic information for object recognition, helping to locate and identify the boundaries and shapes of the target object, and supplementing the parts of the target that are not captured.

[0095] The two sub-models of object recognition and video detection can support and enhance each other to achieve accurate recognition and analysis of abnormal events and important regions in the video: The object recognition model can identify and locate specific target objects in the video, providing information on important regions for the input of the video detection model; The video detection model can use the target information output by the object recognition model, continuously identify the target hit situation in each region of the video frame through the monitoring module. If a key target is found, detect abnormal events related to the target or continuously track and analyze the appearance and subsequent motion trajectory of a specific target. At the same time, according to the real-time monitoring results, use the large model intelligent reasoning to complete the all-round data of the target.

[0096] Pixel semantic segmentation and video detection models can be combined for accurate analysis and recognition of important regions and specific behaviors in videos: The pixel semantic segmentation model can provide detailed pixel-level segmentation results. Based on the specific features of different pixels, the image can be divided into several regions with different sizes and shapes, avoiding the target being split into different image blocks due to standard grid segmentation. The regions hitting specific targets will be regarded as key regions, thus providing more accurate key region information for the video detection model. The video detection model can utilize the semantic segmentation results output by the pixel semantic segmentation model, and respectively view the semantic information involved in different pixel blocks after segmentation. If there are risk events or key targets in the semantic information, it will continuously detect the occurrence and evolution of specific behaviors or abnormal events in the video.

[0097] When the commander interacts with the system in the form of voice, the specific steps of the voice processing model are as Figure 3 shown, including the following steps:

[0098] (1) Voice segmentation and signal preprocessing: Preprocess the input voice signal, including operations such as removing possible background noise, equalizing the volume, removing silent segments, and intelligently filling in missing segments, to improve the quality and clarity of the voice signal.

[0099] (2) Feature extraction: Extract features from the preprocessed voice signal. Commonly used feature extraction methods include the Short-Time Fourier Transform (STFT) or its variants, which convert the voice signal into a spectral representation in the time-frequency domain. Further, Mel-frequency Cepstral Coefficients (MFCC) or other spectral features can be extracted to capture the important feature information of the voice signal and obtain voice features.

[0100] (3) Voiceprint recognition: Before docking with the business system, it is necessary to ensure the user's operation authority first. Therefore, the extracted voice features need to be input into the voiceprint recognition model. The voiceprint recognition model can adopt algorithms such as the Gaussian Mixture Model (GMM) or deep learning models (such as convolutional neural networks or long short-term memory networks) to model and recognize the voice features to achieve voiceprint recognition and verification.

[0101] (4) Speech recognition: For the speech recognition task, the extracted voice features are input into the speech recognition model. The speech recognition model can adopt algorithms such as the Hidden Markov Model (HMM) or deep learning models (such as recurrent neural networks or transcription attention models) to model and recognize the voice features, convert them into corresponding text representations, and dock and input them into the target business system.

[0102] (5) Answer generation and speech synthesis: After the business process is completed, the answer text can be returned. These texts can be converted into voice signals through the speech synthesis module to achieve the function of voice answering. The speech synthesis module can adopt technologies such as synthesis algorithms based on splicing units, parameter generation models, or deep learning models (such as generative adversarial networks) to convert the text into natural and fluent voice signals.

[0103] In the speech processing model, the speech signal processing sub-model is respectively associated with the voiceprint recognition and speech recognition sub-models. As Figure 4 shown, they jointly complete the task of speech processing before the business system docking, and obtain the text content of the user identity and the information input by the user. The speech signal processing sub-model will denoise the original speech data and supplement the speech segments with poor signals or missing segments based on the large model. Then, the effective signals will be converted into feature representations, and the features will be input to the voiceprint recognition model to obtain the user identity and permissions, and input to the speech recognition model to obtain the text sequence of the user input question. Then, the results recognized by the two models can be docked into the business system, and the output result of the business system will call the subsequent speech generation model to convert the answer into a voice signal.

[0104] As Figure 5 shown, the results of video mining belong to the key information recognized, which can be retrieved or viewed by the user. When the user is inconvenient to directly view, the system can describe the key information to the user in voice, and the user can also ask questions to the system in voice and communicate with the system in a dialogue manner.

[0105] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules, modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units, modules or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0106] The unit may or may not be physically separated. The components shown as units can be a physical unit or multiple physical units, that is, they can be located in one place, or they can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0107] In addition, in each embodiment of the present invention, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0108] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), the above-mentioned functions defined in the method of the present invention are executed. It should be noted that the above-mentioned computer-readable medium of the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above.

[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0110] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A multimodal real-time interactive decision-making method, characterized in that: The following steps are involved: Acquire video data and voice signals; Preprocessing the video data to obtain preprocessed video data; Using multimodal generation technology, the image information is supplemented and enhanced through a large model to improve the integrity of the picture. The target detection algorithm is used to detect and locate the target in the preprocessed video data and further classify the targets detected in each preprocessed video data to identify the category and attributes of the target object. According to the location and category information of the target object, based on the deep learning related algorithm, the pixels in the pre-processed video data are semantically segmented, and they are assigned to different semantic categories. The segmentation results are post-processed to obtain pixel-level semantic segmentation results; Extract important areas from pre-processed video data through visual saliency detection technology; Analyze important areas through optical flow estimation method to obtain abnormal motion behavior; The target object, segmentation results and abnormal motion behaviors are integrated to obtain the video mining results; Preprocessing the speech signal to obtain a preprocessed speech signal; Perform feature extraction on the preprocessed speech signal to obtain speech features; Perform voiceprint recognition and speech recognition on the speech features to obtain the question text; The voiceprint recognition is used to obtain user identity and authority; Generate answer text according to question text; Convert the answer text into a voice signal to obtain an answer voice signal; Send the video mining results and answer voice signals to the user; The target recognition model provides the location and category information of the target object, helping the pixel semantic segmentation model to perform pixel-level classification in the target area and perform generative reasoning on the area of ​​the specific target to obtain more information; The pixel semantic segmentation model provides semantic information for target recognition, helps locate and identify the boundaries and shapes of target objects, and supplements the parts of the target that are not captured; The video detection model uses the target information output by the target recognition model to continuously identify the target hits in various areas of the video screen through the monitoring module. If a key target is found, it detects abnormal events related to the target or continuously tracks and analyzes the appearance and subsequent movement trajectory of a specific target. At the same time, based on the real-time monitoring results, the large model intelligent reasoning is used to complete the full range of target data; Pixel semantic segmentation provides important area information for video detection models; The video detection model uses the semantic segmentation results output by the pixel semantic segmentation model to check the semantic information involved in different pixel blocks after segmentation. If the semantic information hits risk events or key targets, it will continue to detect the appearance and evolution of specific behaviors or abnormal events in the video.

2. The multimodal real-time interactive decision-making method according to claim 1, characterized in that: The following steps are also included: The target object is supplemented according to the segmentation results or abnormal motion behaviors.

3. The multimodal real-time interactive decision-making method according to claim 1, characterized in that: The important area is the location of a target object, or an area of ​​a semantic category extracted from the segmentation result.

4. The multimodal real-time interactive decision-making method according to claim 3, characterized in that: The optical flow estimation method is used to analyze important areas and obtain abnormal motion behaviors, which specifically includes the following steps: According to the target object or segmentation result, the motion of the object in the important area is analyzed to obtain abnormal motion behavior.

5. The multimodal real-time interactive decision-making method according to claim 4, characterized in that: The target detection algorithm adopts a region-based convolutional neural network or a single-stage detector; The semantic segmentation network model adopts a convolutional neural network or a fully convolutional neural network; The preprocessing of the video data includes video frame extraction, inter-frame compensation processing and video size unification processing.

6. The multimodal real-time interactive decision-making method according to claim 1, characterized in that: The feature extraction of the preprocessed speech signal is performed to obtain speech features, which specifically includes the following steps: Through short-time Fourier transform or its variants, the speech signal is converted into a spectrum representation in the time-frequency domain, and the Mel spectrum coefficient features are extracted to obtain the speech features.

7. The multimodal real-time interactive decision-making method according to claim 1, characterized in that: The voiceprint recognition adopts a Gaussian mixture model or a deep learning model; The speech recognition adopts a hidden Markov model or a deep learning model; The speech signal conversion adopts a synthesis algorithm based on splicing units, a parameter generation model or a deep learning model; The preprocessing of the speech signal includes background noise, volume equalization, removal of silent segments and intelligent completion of missing segments.

8. A multimodal real-time interactive decision-making system, used to implement the multimodal real-time interactive decision-making method as described in any one of claims 1 to 7, characterized in that: include: An acquisition module, used for acquiring video data and voice signals; A video preprocessing module is used to preprocess the video data to obtain preprocessed video data; The video mining module is used to perform target recognition, semantic segmentation and video detection on the pre-processed video data to obtain the target object, segmentation results and abnormal motion behavior; The fusion module is used to fuse the target object, segmentation results and abnormal motion behavior to obtain the video mining results; A speech preprocessing module is used to preprocess the speech signal to obtain a preprocessed speech signal; A feature extraction module is used to extract features from the preprocessed speech signal to obtain speech features; The voiceprint and speech recognition module is used to perform voiceprint recognition and speech recognition on speech features to obtain the question text; An answer generation module is used to generate an answer text according to the question text; A speech synthesis module is used to convert the answer text into a speech signal to obtain an answer speech signal; The sending module is used to send the video mining results and the answer voice signal to the user.

Citation Information

Patent Citations

  • Artificial intelligence learning method based on speech recognition

    CN109410911A

  • Multi-view-angle-based abnormal behavior analysis system and analysis method for complex places

    CN114708537A

  • Multi-modal man-machine interaction system

    CN117762372A