Data interaction method, electronic device and computer storage medium
By combining a large multimodal model of audio and video streams, generating multimodal structured data and utilizing a multimodal anchor mechanism, the problem of insufficient multimodal data processing in traditional speech recognition technology is solved, efficient and secure multimodal data collection and interaction are achieved, and data interaction efficiency and system reliability are improved.
Patent Information
- Application Number
- CN202510827287.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional speech recognition technology cannot effectively integrate video information for multimodal data processing in audio and video calls, resulting in the system being unable to provide efficient and secure multimodal data collection and real-time interactive experience. There are problems such as missing multimodal information, insufficient audit traceability, and high LLM generative reasoning latency.
By combining a multimodal large model of audio and video streams, user voice descriptions or video images are analyzed in real time to generate multimodal structured data. The multimodal anchor mechanism is used to achieve cross-channel binding and traceability of data, and real-time collaborative processing is carried out in conjunction with the IMS data channel and the large language model.
It realizes the real-time collection and interaction of multimodal data, improves data interaction efficiency and system reliability, ensures data integrity and auditability, and supports various terminals and industry applications in 5G call scenarios.
Smart Images

Figure CN120676203A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of communication technology, and more specifically, to a data interaction method, an electronic device, and a computer storage medium. Background Art
[0002] In the current context of rapid development of informatization and intelligence, in call scenarios (such as audio and video calls), traditional methods of automatically extracting semantics through speech recognition mostly rely on user voice descriptions, converting the content into text, and then extracting key fields for data entry. However, this solution mainly relies on speech recognition and matching rules, processing only audio information, and has obvious shortcomings when dealing with complex or information-intensive communication content:
[0003] (1) Lack of multimodal information: Traditional speech recognition often focuses only on audio information and ignores the rich visual clues in video information, such as user identity confirmation and body language;
[0004] (2) Poor audit traceability: In call services, there is a lack of a full-link recording mechanism, and there is no reliable connection between the recorded data and the call content;
[0005] (3) Lack of timeliness: In real-time business processing scenarios, LLM (Large Language Model) cannot achieve millisecond-level feedback even when using streaming output.
[0006] (4) Difficulty in locating the original data segment: During a call, although structured field data can be output, it is impossible to accurately and quickly locate the specific location in the original call corresponding to this data.
[0007] In summary, no effective solution has been proposed in the related art. Summary of the Invention
[0008] The embodiments of the present application provide a data interaction method, an electronic device, and a computer storage medium to at least solve the problem that traditional automatic form-filling systems based on voice recognition cannot effectively integrate video information for multimodal data processing, resulting in the system being unable to provide efficient and secure multimodal data collection and real-time interactive experience.
[0009] According to one embodiment of the present application, a data interaction method is provided, including: obtaining audio and video data; extracting key information and key frames from the audio and video data; obtaining multimodal structured data based on the key information and the key frames; and performing data interaction through the multimodal structured data.
[0010] According to another embodiment of the present application, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the steps of any one of the above method embodiments when run.
[0011] According to another embodiment of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in the above method embodiment.
[0012] According to another embodiment of the present application, a computer program product is provided, including a computer program, which implements the steps in the above method embodiment when executed by a processor.
[0013] Through the above-mentioned embodiments of the present application, a data interaction method is provided, which can extract key information and key frames from the acquired audio and video data, obtain multimodal structured data based on the key information and key frames, and perform data interaction through the multimodal structured data, that is, it breaks through the limitation of traditional call services that only rely on audio information recognition to generate text, enriches the dimension of data collection, and parses audio information or video information from the acquired audio and video data in real time to obtain more comprehensive multimodal structured data, wherein the multimodal structured data includes not only the text in the audio, but also the text recognized from the video, forming a multi-dimensional structured field list. Therefore, the embodiments of the present application can solve the problem that the traditional automatic form filling system based on voice recognition cannot effectively integrate video information for multimodal data processing, resulting in the system being unable to provide efficient and secure multimodal data collection and real-time interactive experience, thereby achieving the effect of improving data interaction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a hardware structure block diagram of a computer terminal according to the data interaction method of an embodiment of the present application;
[0015] Figure 2 This is a system structure diagram of the operation data interaction method according to an embodiment of the present application;
[0016] Figure 3 This is a schematic diagram of the overall process of the data interaction method according to an embodiment of the present application;
[0017] Figure 4 is a flow chart of a data interaction method according to an embodiment of the present application;
[0018] Figure 5 is a flow chart of structured data collection according to an embodiment of the present application;
[0019] Figure 6is a flowchart of anchor point extraction according to an embodiment of the present application;
[0020] Figure 7 This is a flowchart of real-time semantic extraction according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] The embodiments of the present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0022] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.
[0023] The embodiments of the present application include but are not limited to supporting the following environments: a communication network environment based on a 5G SA architecture, based on VoNR (Voice over New Radio) calls and IMS (IP Multimedia Subsystem) data channel capabilities, terminals with call functions (such as smart phones, remote robots, unmanned customer service terminals, etc.), and business platforms with call functions.
[0024] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a computer terminal as an example, Figure 1 1 is a hardware structure block diagram of a computer terminal according to the data interaction method of an embodiment of the present application. Figure 1 As shown, the computer terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data. The computer terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer terminal. For example, the computer terminal may also include Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0025] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the data interaction method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0026] The transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a communications provider of a computer terminal. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0027] Currently, traditional real-time audio and video extraction and intelligent semantic parsing technologies haven't yet addressed the synchronous conversion of audio and video streams into multimodal structured data containing visual information. For example, user identity authentication and body language (gestures or expressions) recognition haven't yet been performed through video. Field filling relies on speech semantic matching, lacking processing of video content. Furthermore, technologies haven't yet been proposed for directly binding field values to audio and video time segments or video frames. Most focus on speech recognition text alignment, mapping field values back to the audio and video timeline.
[0028] With the development of generative AI (Artificial Intelligence), large language models have demonstrated powerful capabilities in natural language understanding. The parameters of the basic models can range from hundreds of megabytes to billions. The latency difference for the same input text under different inference configurations can reach hundreds of milliseconds to several seconds. Even with streaming output, the system needs to wait for the large language model to output multiple tokens cumulatively before determining the field value, making it impossible to output field values at the millisecond level.
[0029] In summary, the traditional technology still has the following deficiencies:
[0030] (1) Unable to locate the original audio and video clips: Traditional technologies only output structured field data, but do not record the input voice or time position corresponding to the structured field data. Therefore, it is impossible to quickly trace back to the original audio and video clips during auditing.
[0031] (2) Lack of multimodal information: Traditional speech recognition often only focuses on audio information and ignores visual clues in the video (such as user identity authentication, body language, etc.), making it impossible to output multimodal information.
[0032] (3) Insufficient audit traceability: In call services, there is a lack of transparent full-link records and a lack of reliable connection between the recorded data and the audio and video content, making it difficult to meet the requirements during the audit process.
[0033] (4) The latency of LLM generative reasoning is high: Even with streaming output, it is difficult to achieve millisecond-level feedback. This makes it impossible to rely on large language models to extract field values instantly in real-time systems, and the system response lacks timeliness.
[0034] Therefore, in order to improve the accuracy of data entry and system reliability, a new solution is urgently needed to synchronously process audio and video streams, extract semantic fields in real time, and establish an auditable traceability mechanism.
[0035] In view of the above problems, this application embodiment proposes a real-time structured data collection and interaction method based on call services and multimodal large models. For example, in VoNR (Voice over New Radio) voice or video calls, it is no longer limited to relying on audio information recognition to generate text. It can use streaming voice recognition, video OCR (Optical Character Recognition) and liveness detection technology to parse user voice descriptions or video images in real time and convert them into structured data. The structured data includes not only text, but also video information in the video (such as OCR-recognized text, liveness detection results, etc.), forming a multi-dimensional structured field.
[0036] In addition, the user provider side sends an editable "data card" to the terminals of both parties through the IMS Data Channel (IMSDC). The content of the "data card" can be defined according to the business template, and dynamic filling and updating are supported, realizing the transition from single-modal data to multi-modal data, significantly improving the comprehensiveness and practicality of data collection.
[0037] In order to meet the audit requirements of high-compliance industries for "searchability and locatability", the embodiment of this application also proposes a "multimodal anchor" mechanism, that is, by synchronously carrying absolute timestamps in the audio and video streams, the voice segments and video key frames near the same time are chain-bound to form a cross-channel unified, traceable and non-repudiation anchor. This "multimodal anchor" mechanism ensures that each structured field can be accurately located to a certain moment in the call to facilitate subsequent compliance audits and backtracking. For example, the identity information or signature screen provided by the user during the call can be quickly located and its authenticity verified through the anchor point.
[0038] Through the above embodiment, technologies such as VoNR calls, IMS data channels, and large language model semantic understanding are integrated to achieve real-time data collection and interaction. At the same time, a "real-time collaboration mechanism" is provided. After the field is updated, it can be immediately encapsulated into a data message (event), which is pushed safely and reliably to the user's (such as medical, industrial, etc.) backend system through the user provider's message queue. After processing, the field update results can be fed back to the user, realizing "chatting and doing". This real-time collaboration mechanism not only improves the user experience, but also improves business processing efficiency.
[0039] In summary, the embodiments of the present application provide a new solution for data collection and interaction in 5G (5th Generation Standalone) call scenarios through multimodal structured data, multimodal anchor mechanism and real-time collaboration mechanism, which can support a variety of terminals and industry applications to meet the needs of different scenarios.
[0040] Figure 2 is a system structure diagram of the operation data interaction method according to an embodiment of the present application, such as Figure 2 As shown, for example, in a VoNR voice or video call, the data interaction system includes: a new call management module, a new call media processing module, a new call service processing module, a new call signaling processing module, and a structured data processing module.
[0041] New call management module: responsible for the registration / deregistration of new call users, the distribution / deletion of materials, the customized settings of users' new calls, and the related settings of the new call system. It also serves as a bridge for access to third-party systems, allowing external triggering of new call services.
[0042] New call media processing module: Responsible for establishing and setting up media and data channels, it has media interface processing capabilities. It can also call ASR (Automatic Speech Recognition) to convert speech into text, call NLU (Natural Language Understanding) to identify keywords in speech, and call LLM (Large Language Model) for multimodal data recognition.
[0043] New call service processing module: responsible for receiving requests from the new call management module and the new call signaling module, completing the logic of the service control plane according to the corresponding process and database configuration, and receiving requests issued by the third-party system through the new call management module to complete the corresponding service processing.
[0044] New call signaling processing module: Responsible for processing SIP (Session Initiation Protocol) signaling between the caller and the called party, reporting user-level subscription call events and call parameters, and completing the creation, update, and deletion of the media plane according to the control logic of the new call service processing module.
[0045] Structured data processing module: Responsible for extracting keywords based on the content reported by the media and then calling the structured field model. It supports field structure generation and anchor binding (such as fields, start and end times, text fragments, and other information), supports two-way synchronization of fields between the user end and third-party systems, supports field status change prompts, conflict handling (such as time and role priority), reception confirmation status, and execution through the new call service module.
[0046] Through the above-mentioned data interaction system, the media channel, namely the audio / video SRTP (Secure Real-time Transport Protocol) and IMS-DC can be bound to the same SDP (Session Description Protocol). After the session starts, voice is anticipated while multimodal structured data is pushed in real time. At the same time, the read-only "data card" in the DC is upgraded to a writable structured data stream. Combined with ASR and event network management, all elements described by the user in the call or "shown to the camera" can be converted into fields and anchors and sent synchronously to the third-party system. Therefore, real-time data transmission and interaction are achieved, while ensuring data traceability and non-tampering.
[0047] Figure 3 This is a schematic diagram of the overall process of the data interaction method according to an embodiment of the present application. Figure 3 As shown, for example, in a VoNR voice or video call, the overall process of the data interaction method includes:
[0048] Step 1: When a user terminal initiates a new call, it requests the call signaling processing module to establish the call. After receiving the request, the call signaling processing module parses the signaling of the call process and reports the call connection event to the call service processing module.
[0049] Step 2: The call service processing module determines whether the third-party system has registered a new call service;
[0050] If a third-party system registers a new call service, the call service processing module activates the anchor binding function and anchors the audio and video streams through the "multimodal anchor" mechanism;
[0051] Furthermore, by synchronously carrying absolute timestamps in the audio and video streams, and chain-binding the voice segments and video key frames near the same moment, a data chain on a unified timeline is formed, which can ensure that each structured field can correspond to a certain moment in the call process. This not only greatly improves the integrity and reliability of the data, but also improves the security and accuracy of the audit, and prevents misjudgments due to data loss or unclear records.
[0052] Step 3: Establish data channels between the user terminal and the call media processing module, and between the call media processing module and the third-party system;
[0053] Step 4: The call media processing module returns the anchoring result to the call signaling processing module, which then transmits the anchoring result to the call service processing module.
[0054] Step 5: After the anchoring is successful, the call service processing module sends a structured data processing request to the call media processing module;
[0055] Step 6: The call media processing module reports the real-time video recognition results of ASR / NLU / LLM to the structured data processing module in real time;
[0056] Step 7: The structured data processing module performs corresponding processing based on the recognition results, searching for a preset template or generating a form based on the user's voice description. Specifically, during the communication process, the intelligent processing module fills out the form or extracts the anchor point record based on the recognition results, and then reports the obtained multimodal structured data to the call media processing module.
[0057] Step 8: After both parties confirm the communication, the call media processing module reports the multimodal structured data to the user terminal and the third party through the data channel;
[0058] Step 9: After obtaining the multimodal structured data, the third party conducts business audit and processing, and sends the final processing results to the call management module;
[0059] Step 10: The call management module sends the final processing result to the call service processing module for display, and then sends it to the call media processing module, which displays the final processing result to the user terminal and the third party.
[0060] Through the above embodiment, the multimodal anchor mechanism, multimodal structured data and LLM-NLU dual-stream parsing are combined to achieve high-precision real-time data positioning and integrity verification. At the same time, it not only greatly improves the immediacy and accuracy of business processing, but also provides convenience for subsequent audits. In particular, when service providers handle customer business, they can achieve efficient and secure data collection and interaction through new call technology to ensure the auditability of business processes.
[0061] Figure 4 is a flow chart of a data interaction method according to an embodiment of the present application. Figure 4 As shown, the data interaction process includes the following steps:
[0062] Step S402: Acquire audio and video data;
[0063] In this embodiment, the acquired audio and video data may include valid audio data and / or valid video data.
[0064] In some embodiments, before obtaining the audio and video data, the method also includes: generating a field template based on a pre-configured prompt word or a third-party system voice, and issuing an interactive language based on the field template, wherein the interactive language is sent to the user terminal through a data channel so that the user can perform corresponding operations according to the interactive language.
[0065] Step S404: extracting key information and key frames from the audio and video data;
[0066] In some embodiments, extracting key information and key frames from the audio and video data includes: parsing the audio and video data through automatic speech recognition and natural language understanding to extract valid audio data; cutting out video frames in the audio and video data through manual keystrokes, voice commands or target detection to extract valid video data and key frames.
[0067] In this embodiment, Figure 5 is a flow chart of structured data collection according to an embodiment of the present application, such as Figure 5 As shown, in a new call, you need to perform the following steps first:
[0068] S5-1, load field template;
[0069] Load field templates based on pre-configured prompt words (such as accident, license plate number, etc.), or load field templates based on customer service voice description keywords from a third-party system;
[0070] For example, when the system detects the keyword "accident", it will load a template containing fields such as accident time, location, accident description, and vehicle damage information; when the system detects the prompt word "VIN" or "license plate number", it will trigger the loading of fields related to vehicle identity information.
[0071] S5-2, issue guiding language (interactive language);
[0072] The customer service of the third-party system generates a guide message according to the field template. The guide message may include text, voice, AR (Augmented Reality) viewfinder, etc., and is sent in real time through the data channel (DC channel).
[0073] S5-3. Extract voice from the audio and video data and parse it through streaming ASR (automatic speech recognition) and NLU (natural language understanding) to obtain text fields.
[0074] S5-4, video data acquisition method;
[0075] (1) Manual keystroke: The administrator or third-party customer service clicks a keystroke to trigger a capture and save the keyframe;
[0076] (2) Voice command: Real-time NLU recognition (for example, recognizing keywords such as "I took a picture of the rear of the car"), cutting frames, and saving keyframes;
[0077] (3) Object detection: A real-time object detection algorithm such as lightweight YOLO (object detection algorithm) is used to detect a VIN (Vehicle Identification Number), license plate, or damage frame in the viewfinder, and the frame is captured and the keyframe is saved.
[0078] S5-5: During the call, the video field is extracted through OCR and liveness detection, and the text field and the video field are merged. S5-4 can be executed continuously if the call is not ended.
[0079] S5-6, Draft Upload: The voice field and the visual field related information (including anchors) are combined into a JSON (JavaScript Object Notation) and uploaded to the draft table through the data channel.
[0080] In some embodiments, after obtaining the audio and video data, the method further includes: mapping the relative time of the audio and video data to absolute time; obtaining key information and key frames of the audio and video data after mapping to the absolute time, encapsulating the key information and processed key frames to obtain a multimodal field anchor.
[0081] In this embodiment, a multimodal field anchor needs to be generated after obtaining the audio and video data. For example, in an audio and video call, the audio and video data is transmitted in a relative time format. In order to facilitate auditing and backtracking, the relative time of the audio and video data is mapped to absolute time. On the basis of ensuring that the audio and video data is synchronized with the absolute time axis, key information and key frames are extracted from the audio and video data, and the extracted key information and key frames are merged and encapsulated into structured data of a multimodal field anchor. The multimodal field anchor is transmitted between the two parties of the call through a DC channel, ensuring real-time sharing and recording of data, which facilitates subsequent data auditing and backtracking.
[0082] In some embodiments, mapping the relative time of the audio and video data to the absolute time includes: when a real-time transport protocol data packet is received and the real-time transport protocol data packet carries an absolute capture timestamp, time-pairing the absolute capture timestamp of the audio and video data with the network time protocol to obtain the absolute time; or, when the real-time transport protocol data packet does not carry the absolute capture timestamp and a real-time transport control protocol sender report is received, recording the real-time transport protocol timestamp of the audio and video data carried in the sender report and the coordinated universal time on the same line to obtain the absolute time; or, when the real-time transport protocol data packet does not carry the absolute capture timestamp, no real-time transport control protocol sender report is received, and a data channel comparison message is received, mapping the real-time transport protocol timestamp of the audio and video data to the coordinated universal time according to the data channel comparison message and the current interpolation to obtain the absolute time.
[0083] In this embodiment, mapping the relative time of audio and video data to absolute time includes but is not limited to the following three methods: absolute time mapping based on ACT (Absolute-Capture-Time), absolute time mapping based on RTCP (Real-time Transport Control Protocol) SR (Sender Report), and absolute time mapping based on DC.
[0084] Furthermore, when receiving an RTP (Real-time Transport Protocol) packet, it is determined whether the RTP packet carries an ACT header extension field. If it carries an ACT header extension field, absolute time mapping is performed based on the timestamp of the ACT header extension field; if it does not carry an ACT header extension field, absolute time mapping is performed by parsing the RTCP SR; if it does not carry an ACT header extension field and the RTCP SR cannot be parsed, absolute time mapping is performed through the DC control line.
[0085] In this embodiment, Figure 6 is a flowchart of anchor point extraction according to an embodiment of the present application, such as Figure 6 As shown, mapping the relative time of audio and video data to absolute time includes the following process:
[0086] (1) Prioritize reading the ACT header extension field in the RTP packet. RTP is a protocol used to transmit audio and video data on the network, ensuring that data can be transmitted quickly and in sequence. RTP not only transmits the audio and video content itself, but also provides timestamps and sequence numbers for the audio and video content to facilitate the receiver to reassemble and synchronize the data.
[0087] If the RTP packet carries ACT, the timestamp is immediately matched with the NTP (Network Time Protocol) time. The timestamp helps the receiving end synchronize the time of the audio and video data to a unified time standard, which can solve the problem of inconsistent timestamps in the audio and video streams.
[0088] (2) If the RTP packet does not carry ACT, the time calibration is performed based on the RTP timestamp in the RTCP SR and the corresponding UTC time (Coordinated Universal Time). The Sender Report is a report sent periodically by RTCP. The Sender Report contains RTP statistical information, such as the current RTP timestamp and the corresponding UTC time. The receiver can calibrate the RTP timestamp based on the Sender Report.
[0089] In this embodiment, each time the receiver receives an SR, it will update the "RTP-UTC absolute clock" comparison line in combination with the time mapping algorithm. That is, the RTP timestamp of a certain moment and the UTC time of that moment are placed in the same record line, and the absolute time of the audio and video stream is calculated in real time to ensure the timing consistency of the audio and video.
[0090] (3) If RTCP cannot be relied upon, obtain the data channel control line;
[0091] When the RTCP Sender Report cannot arrive, a (rtp_ts,utc_us) comparison message is sent every few seconds (for example, every 5 seconds) through the data channel. The receiver can only retain the most recent message and map the subsequent RTP packet timestamps to the unified UTC time based on the most recent message using the current interpolation method.
[0092] Through the cooperation of the extended field ACT of the RTP packet and the RTCP sender report in the above embodiment, time synchronization of audio and video streams is achieved, so that multiple data sources (for example, audio, video, images, etc.) can be processed based on a unified absolute time standard, thereby ensuring the timing consistency and accuracy of multimodal structured data. In addition, this embodiment does not require modification of the terminal codec. For example, if the call service requires high precision, the sending period can be reduced to 1 second. The implementation process remains unchanged, which can meet the needs of most scenarios.
[0093] In some embodiments, the key information and key frames of the audio and video data mapped to the absolute time are obtained, and the key information and the processed key frames are encapsulated to obtain a multimodal field anchor, including: obtaining a key information key-value pair and a key information start and end timestamp; removing the header information in the key frame and obtaining a payload, and processing the payload through a hash function to obtain a key frame fingerprint; encapsulating the key information key-value pair, the key information start and end timestamp and the key frame fingerprint to obtain a multimodal field anchor.
[0094] In this embodiment, after the relative time of the audio and video data is mapped to the absolute time, the key information key-value pair and the key information start and end timestamps are obtained, such as Figure 6 As shown, for audio information, the start and end timestamps of text fields (key information) are captured through streaming ASR (automatic speech recognition) and NLU (natural language understanding); for video information, key frames are captured through OCR and liveness detection technology.
[0095] Locate a keyframe (IDR for H.264 / H.265) near the end of a segment on the audio / video timeline, but no later than that. Remove the header information from the keyframe to obtain the video payload. Apply a cryptographic hash function (e.g., SHA-256) to the payload and obtain the keyframe fingerprint `keyframe_hash`. If the current group contains no IDR (Independent Decoding Refresh) frames, temporarily obtain the fingerprint of the most recent P-frame and replace it when an IDR frame is encountered. Keyframe fingerprints can be used to detect clips or replacements during audits.
[0096] After obtaining the keyframe fingerprint, the event scheduler serializes {field_key, field_value, ntp_start_us, ntp_end_us, keyframe_hash, confidence} into a JSON message (encapsulating the field key-value pair, start and end timestamps, keyframe fingerprint, and other metadata (such as confidence) into a JSON message), that is, encapsulating it into a multimodal field anchor and transmitting the multimodal field anchor through the data channel.
[0097] In some embodiments, after encapsulating the key information key-value pair, the key information start and end timestamps, and the key frame fingerprint to obtain a multimodal field anchor, the method further includes: obtaining a segment chain of the audio and video data through a data window and a Merkle tree algorithm, and writing the segment chain and the multimodal field anchor into a data log together, wherein the data log carries a digital signature.
[0098] In this embodiment, if Figure 6 As shown in the example, using a hash algorithm, after obtaining a multimodal field anchor, the system can splice the audio and video bitstreams of the same segment in a two-second window and calculate the segment hash value 'segment_hash'. It then uses a Merkle-Tree algorithm to iteratively merge these segments to obtain a root hash 'root_hash' (segment hash chain), which is then uploaded via the data channel. The root hash 'root_hash' and the multimodal field anchor are simultaneously written to a digitally signed audit log. This process ensures unique data representation, ensuring that even minor changes will result in a significant change in the segment hash value, effectively preventing audio and video data tampering. Subsequent audits can directly play the audio and video based on the multimodal field anchor. Combining the keyframe_hash and the corresponding Merkle path enables millisecond-level location and traceability, eliminating the need for time-consuming recalculation and searching for corresponding information. This not only speeds up the audit process but also improves audit accuracy, preventing misjudgments due to data loss or unclear records.
[0099] In some cases, under poor network conditions, the system may not be able to obtain the ACT field in the RTP packet, or RTCP-SR transmission may be blocked. In this case, the system will use an alternative solution to periodically send timestamp comparison lines over the data channel to maintain data synchronization and time accuracy. The frequency and format of timestamp comparison lines can be remotely adjusted according to business needs without restarting the session.
[0100] If the network conditions are restored, the system will automatically switch back to using the ACT field or RTCP-SR in the RTP packet to obtain and synchronize timestamps, without affecting the positioning and recording of multimodal fields.
[0101] Step S406: Acquire multimodal structured data according to the key information and the key frame;
[0102] In this embodiment, in a real-time new call scenario, users' voice conversations and video interactions simultaneously generate a large amount of data, which contains rich business information. Speech recognition technologies (such as ASR and NLU) extract key information in text form from the voice stream, and video analysis technologies (such as OCR and liveness detection) identify and capture key frames from the video stream. These include, but are not limited to, visual clues such as the user's gestures, facial expressions, and presented IDs or documents.
[0103] In this embodiment, multimodal structured data is a data type that converts data from multiple sensory inputs (such as voice and video) into a unified structured format. This data type is not limited to text information, but rather converts different modalities such as text, images, and video into a processable format.
[0104] For example, multimodal structured data mainly includes the following two aspects:
[0105] Voice field information: text information extracted from voice streams using voice recognition technology, such as users' spoken commands, account information, transaction details, etc.
[0106] Visual field information: key frames identified from the video stream, including but not limited to text on paper documents, displayed documents, facial expressions, gestures, and other non-verbal communication signals identified through video analysis technology.
[0107] In some embodiments, after obtaining the multimodal structured data based on the key information and the key frame, the method further includes: extracting fields in the multimodal structured data; and integrating the fields into a multimodal structured data container based on the fields and field confidence.
[0108] In this embodiment, fields are extracted from the acquired multimodal structured data, and each field is assigned a field confidence. The confidence evaluation mechanism can ensure the reliability of the field and prevent low-quality data from being incorrectly integrated into the multimodal structured data container, where the multimodal structured data container includes but is not limited to data forms, databases, etc., which are used to store field data of different modalities.
[0109] In some embodiments, integrating the field into the multimodal structured data container based on the field and the field confidence includes: obtaining a target confidence of the field based on a first weight of the average word-level probability of the field and a second weight of the field confidence; integrating the field into the multimodal structured data container when the target confidence is greater than or equal to a first preset threshold; or, integrating the field into the multimodal structured data container through user confirmation when the target confidence is greater than or equal to a second preset threshold and less than the first preset threshold; or, integrating the field into the multimodal structured data container through manual review and joint confirmation of a large language model when the target confidence is less than the second preset threshold.
[0110] In this embodiment, Figure 7 is a flowchart of real-time semantic extraction according to an embodiment of the present application, such as Figure 7 As shown, based on the average word-level probability of ASR and the corresponding field confidence of NLU, the two are used to calculate the total confidence according to certain weights. For example, the total confidence (target confidence) is calculated according to the first weight of the average word-level probability of ASR and the second weight of the field confidence. The first weight and the second weight can be configured in the call management module.
[0111] Different processing decisions are made based on different target confidence levels (such as low, medium, and high). Each level of target confidence determines a different processing flow, as follows:
[0112] (1) If the target confidence is greater than or equal to a first preset threshold (for example, the first preset threshold is 0.8, which can also be called high confidence), the field can be directly filled in the form;
[0113] (2) If the target confidence is greater than or equal to a second preset threshold (e.g., the second preset threshold is 0.6, which may also be referred to as a low confidence), and is less than the first preset threshold, automatic clarification is performed, i.e., the field is filled in the form through user confirmation;
[0114] (3) If the target confidence is less than the second preset threshold, it is marked as manual review, and manual review and the large language model are used to jointly confirm whether the field is filled in the form.
[0115] In this embodiment, the first preset threshold and the second preset threshold can be set according to actual conditions.
[0116] In some embodiments, when the target confidence is less than the second preset threshold, the field is integrated into the multimodal structured data container through manual review and large language model confirmation, including: verifying the field through a large language model to generate a verification confidence; when the verification confidence is greater than or equal to the first preset threshold, integrating the field into the multimodal structured data container; or, when the verification confidence is greater than or equal to the second preset threshold and less than the first preset threshold, generating a correction field and integrating the correction field into the multimodal structured data container through user confirmation; or, when the verification confidence is less than the second preset threshold, integrating the field into the multimodal structured data container through manual review.
[0117] In this embodiment, if Figure 7 As shown, if the target confidence is less than the second preset threshold, manual review and the large language model are used to jointly confirm whether to fill in the field in the form, including the following process:
[0118] (1) If the target confidence is less than a second preset threshold, a manual review is triggered;
[0119] In this embodiment, the system submits the field and related context (voice clip, video frame, draft value) to manual review.
[0120] (2) Run LLM and manual verification in parallel in the background;
[0121] In this embodiment, in parallel with manual review, a large language model is asynchronously called to perform consistency verification on the field and generate a verification confidence level p_consist.
[0122] (3) Evaluate the verification results of large language models running in parallel with manual review;
[0123] If the verification confidence is greater than or equal to the first preset threshold, the system generates a patch JSON, where the patch JSON only contains the corrected field values, and the corrected field values can be directly filled in the form;
[0124] If the verification confidence is greater than or equal to the second preset threshold and less than the first preset threshold, the system generates a patch JSON containing only the corrected field values, and the user confirms the corrected field values and enters them into the form.
[0125] When the verification confidence is less than the second preset threshold, the large language model is no longer called through manual review, and the corrected field value is filled in the form.
[0126] In some embodiments, the draft form generated by the system and the patch data generated by the large language model are uploaded to the new call service processing module and the third-party system through a pre-established data channel. In order to ensure the security and reliability of data transmission, the data channel can use encrypted communication to prevent the data from being intercepted or tampered with during transmission.
[0127] In this embodiment, anchor points (field anchor points and keyframe fingerprints) are generated for each data upload. The upload is not completed in one go, but is carried out gradually as the call progresses and the data is generated. A complete traceable log is formed through the segment hash chain, which facilitates rapid backtracking when problems arise.
[0128] Step S408: Perform data interaction through the multimodal structured data.
[0129] In this embodiment, multimodal structured data that combines voice and visual information enables more efficient and secure real-time communication service processing. In traditional communication scenarios, data interaction is often limited to a single modality, such as obtaining user information solely through voice recognition. However, data interaction through multimodal structured data in this embodiment greatly expands the breadth of data interaction.
[0130] In some embodiments, the data interaction method further includes: establishing a data channel for transmitting the multimodal structured data.
[0131] In this embodiment, when establishing a call session between a user terminal and a third-party system, the system also needs to initialize the IMS-DC (IMS-data channel) capability. The established data channel not only supports text transmission, but can also be used for real-time transmission of multimodal structured data. Among them, the calls established between the user terminal and the third-party system include but are not limited to VoNR calls, traditional voice calls, and VoLTE (Voice over Long-Term Evolution) calls based on IP (Internet Protocol).
[0132] Through the above-mentioned embodiments of the present application, a data interaction method is provided to obtain audio and video data, extract key information and key frames from the audio and video data, obtain multimodal structured data based on the key information and key frames, and perform data interaction through the multimodal structured data, that is, to break through the limitation of traditional call services that only rely on audio information recognition to generate text, and propose to parse the user's voice description or video picture from the obtained audio and video data in real time and convert it into multimodal structured data, wherein the multimodal structured data includes not only the text in the audio, but also the text recognized from the video, forming a multi-dimensional structured field list. Therefore, the embodiments of the present application can solve the problem that the traditional automatic form filling system based on voice recognition cannot effectively integrate video information for multimodal data processing, resulting in the system being unable to provide efficient and secure multimodal data collection and real-time interactive experience, thereby achieving the effect of improving data interaction efficiency.
[0133] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0134] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0135] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above method embodiments when run.
[0136] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0137] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0138] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0139] According to yet another embodiment of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the method described in each embodiment of the present disclosure are implemented.
[0140] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.
[0141] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed across a network composed of multiple computing devices, they can be implemented using program code executable by the computing device, and thus, they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be performed in a different order than herein, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0142] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A data interaction method, characterized in that: include: Get audio and video data; Extract key information and key frames from the audio and video data; Acquiring multimodal structured data according to the key information and the key frames; Data interaction is performed through the multimodal structured data.
2. The method according to claim 1, characterized in that The method further comprises: A data channel for transmitting the multimodal structured data is established.
3. The method according to claim 1, characterized in that Extracting key information and key frames from the audio and video data includes: Parsing the audio and video data through automatic speech recognition and natural language understanding to extract valid audio data; The video frames in the audio and video data are cut out by manual keystrokes, voice commands or target detection to extract valid video data and key frames.
4. The method according to claim 1, wherein After obtaining the audio and video data, the method further includes: Mapping the relative time of the audio and video data to absolute time; The key information and key frames of the audio and video data mapped to the absolute time are obtained, and the key information and the processed key frames are encapsulated to obtain a multimodal field anchor.
5. The method according to claim 4, characterized in that Mapping the relative time of the audio and video data to absolute time includes: Upon receiving a real-time transport protocol data packet, and the real-time transport protocol data packet carries an absolute capture timestamp, time-matching the absolute capture timestamp of the audio and video data with the network time protocol to obtain an absolute time; or If the real-time transport protocol data packet does not carry an absolute capture timestamp and a real-time transport control protocol sender report is received, the real-time transport protocol timestamp of the audio and video data carried in the sender report and the coordinated universal time are recorded in the same line to obtain the absolute time; or When the real-time transport protocol data packet does not carry an absolute capture timestamp, no real-time transport control protocol sender report is received, and a data channel comparison message is received, the real-time transport protocol timestamp of the audio and video data is mapped to Coordinated Universal Time based on the data channel comparison message and the current interpolation to obtain an absolute time.
6. The method according to claim 4, characterized in that The acquiring of key information and key frames of the audio and video data mapped to the absolute time, and encapsulating the key information and the processed key frames to obtain a multimodal field anchor point includes: Get key information key-value pairs and key information start and end timestamps; Removing header information from the key frame and obtaining a payload, and processing the payload using a hash function to obtain a key frame fingerprint; The key information key-value pair, the key information start and end timestamps, and the key frame fingerprint are encapsulated to obtain a multimodal field anchor point.
7. The method according to claim 6, characterized in that After encapsulating the key information key-value pair, the key information start and end timestamps, and the key frame fingerprint to obtain a multimodal field anchor point, the method further includes: A segment chain of the audio and video data is obtained through a data window and a Merkle tree algorithm, and the segment chain and the multimodal field anchor are written into a data log together, wherein the data log carries a digital signature.
8. The method according to claim 1, characterized in that Before obtaining the audio and video data, the method further includes: A field template is generated according to a pre-configured prompt word or a third-party system voice, and an interactive language is issued according to the field template. The interactive language is sent to the user terminal through a data channel so that the user can perform corresponding operations according to the interactive language.
9. The method according to claim 1, characterized in that After obtaining the multimodal structured data according to the key information and the key frame, the method further includes: Extracting fields from the multimodal structured data; The fields are integrated into a multimodal structured data container according to the fields and the field confidences.
10. The method according to claim 9, characterized in that Integrating the fields into a multimodal structured data container according to the fields and the field confidences includes: Obtaining a target confidence of the field according to a first weight of the average word-level probability of the field and a second weight of the field confidence; In a case where the target confidence is greater than or equal to a first preset threshold, integrating the field into a multimodal structured data container; or, In a case where the target confidence is greater than or equal to the second preset threshold and less than the first preset threshold, integrating the field into the multimodal structured data container through user confirmation; or When the target confidence is less than the second preset threshold, the field is integrated into the multimodal structured data container through manual review and confirmation by a large language model.
11. The method according to claim 10, characterized in that In the case where the target confidence is less than the second preset threshold, integrating the field into the multimodal structured data container through manual review and confirmation by a large language model includes: Verify the field using a large language model to generate verification confidence; In a case where the verification confidence is greater than or equal to the first preset threshold, integrating the field into a multimodal structured data container; or, If the verification confidence is greater than or equal to the second preset threshold and less than the first preset threshold, generate a correction field and integrate the correction field into the multimodal structured data container through user confirmation; or When the verification confidence is less than the second preset threshold, the field is integrated into the multimodal structured data container through manual review.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 11 are implemented.
13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method described in any one of claims 1 to 11 are implemented.
14. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 11 are implemented.