Intelligent conference recording and recording method and system

Through the edge-cloud collaborative architecture, edge devices generate multimodal information and send it to cloud service devices for in-depth analysis, solving the low efficiency and poor quality problems of traditional meeting record and recording methods, and realizing efficient and secure intelligent meeting record and recording.

CN120568005BActive Publication Date: 2025-09-26XIAMEN RGBLINK SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511068063.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-09-26
Estimated Expiration
2045-07-31

AI Technical Summary

Technical Problem

Traditional conference record-keeping and recording methods are time-consuming and labor-intensive, with incomplete and inaccurate records. The recording equipment is cumbersome to operate, the recording quality is low, and it is difficult to achieve efficient integration and clear recording. It also lacks intelligent analysis and processing capabilities, which affects the efficiency of conference material utilization.

Method used

Adopting an edge-cloud collaborative architecture, it obtains audio and video information through edge devices, generates multimodal information, combines timeline processing and sends it to cloud service devices for in-depth analysis and encryption, generates structured meeting records, and supports multi-platform distribution and cross-conference retrieval.

Benefits of technology

It achieves efficient and secure meeting record and recording, supports access to mainstream video conferencing software, ensures data real-time and integrity, and provides a complete intelligent meeting record and recording technology system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120568005B_ABST
    Figure CN120568005B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent conference record and recording method and system. The present invention receives multimodal information sent by an edge device by a cloud service device; processes the multimodal information, combines context-aware error correction, content segmentation, dynamic segmentation strategy and preset storage strategy to generate conference record information; encrypts the conference record information based on participant role information to generate target format conference information; processes the target format conference information to generate initial conference scene mode information including a conference outline, PPT, and mind map; processes the target format conference information and the initial conference scene mode information to generate conference information data after priority sorting and strategy processing; performs semantic analysis, feature fusion and cross-modal association processing on the conference information data to generate target conference scene mode information; confirms the target application based on the target conference scene mode information, and sends target format instruction information to the target application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal data processing, and in particular to an intelligent conference record and recording method and system. Background Art

[0002] In the modern office environment, the demand for intelligent meeting recording and recording systems is increasingly urgent. Traditional meeting recording and recording methods have numerous limitations and cannot meet the demands of efficient office work. Regarding meeting recording, manual recording is not only time-consuming and labor-intensive, but also prone to missing key information such as important decision details, division of responsibilities, and timelines due to the limited attention of the recorder. In some meetings with intense and fast-paced discussions, recorders struggle to keep up with the pace of speech, resulting in incomplete and inaccurate records. Even with the assistance of recording equipment, extracting key information from lengthy recordings and documenting them requires considerable time and effort. While some speech-to-text technologies are currently used for meeting recording, they commonly suffer from inaccurate recognition of specialized terminology. Meetings across different industries often contain a large amount of specialized vocabulary. Common transcription software, lacking in-depth optimization for specific industry terminology, often makes misrecognition errors, seriously impacting the accuracy and usability of the records. Furthermore, the text generated by existing transcription tools often lacks structured processing; it is simply a collection of words, without categorization by meeting topic, discussion items, resolutions, and so on. This makes subsequent review and use of meeting records extremely inconvenient.

[0003] Traditional recording equipment and methods also face numerous challenges in the field of conference recording. At the meeting site, multiple devices must be manually operated to ensure comprehensive capture of both visuals and audio. This is not only cumbersome but can also easily lead to important scenes or sounds being missed due to negligence. For example, in large conference rooms, improper camera angle adjustment may not capture all participants and presentations. Improper microphone placement can result in unclear audio capture and weak sound in certain areas. Furthermore, for remote meetings, existing recording methods struggle to efficiently integrate and clearly record audio and video from different participants. Audio delays, image freezes, and synchronization issues are common, severely impacting the quality and integrity of the conference recordings. Traditional conference recordings often lack intelligent analysis and processing capabilities, unable to automatically identify key content, key speakers, and key decision-making components of the meeting. This makes subsequent retrieval of useful information from a large volume of recorded video like searching for a needle in a haystack, significantly reducing the efficiency of meeting recordings. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0005] An intelligent conference record and recording method, applied to edge devices, includes: obtaining audio information and video information; processing the audio information and video information in combination with a timeline to generate multimodal information, the multimodal information including text information, voice information and image information; sending the multimodal information to a cloud service device for the cloud service device to generate target conference scene information; matching the target hard disk based on preset matching rules, and if the match is successful, sending the multimodal information to the target hard disk for the target hard disk to store the text information, voice information and image information.

[0006] An intelligent conference record and recording method, applied to a cloud service device, comprises: receiving multimodal information sent by an edge device, the multimodal information including text information, voice information and image information; processing the multimodal information, and generating conference record information by combining context-aware error correction, content segmentation, dynamic segmentation strategy and preset storage strategy; encrypting the conference record information based on participant role information to generate target format conference information; processing the target format conference information to generate initial conference scene mode information including a conference outline, PPT and mind map; processing the target format conference information and the initial conference scene mode information based on a priority processing strategy for multimodal information to generate conference information data after priority sorting and strategy processing; performing semantic analysis, feature fusion and cross-modal association processing on the conference information data to generate target conference scene mode information; confirming a target application based on the target conference scene mode information, and sending target format instruction information to the target application for processing by the target application; locating similar decision fragments in historical meetings through semantic search technology, and providing data support and processing capabilities for cross-conference retrieval.

[0007] An intelligent conference record and recording system includes: obtaining audio information and video information by an edge device; processing the audio information and video information in combination with a timeline to generate multimodal information, the multimodal information including text information, voice information and image information; sending the multimodal information to a cloud service device for the cloud service device to generate target conference scene information; receiving the multimodal information sent by the edge device by the cloud service device, the multimodal information including text information, voice information and image information; processing the multimodal information, combining context-aware error correction, content segmentation, dynamic segmentation strategy and preset storage strategy to generate conference record information; encrypting the conference record information based on participant role information to generate target format conference information; and Processing is performed to generate initial conference scene mode information including a conference outline, PPT, and mind map; the target format conference information and the initial conference scene mode information are processed based on the priority processing strategy of multimodal information to generate conference information data after priority sorting and strategy processing; semantic analysis, feature fusion, and cross-modal association processing are performed on the conference information data to generate target conference scene mode information; the target application is confirmed based on the target conference scene mode information, and the target format instruction information is sent to the target application for processing by the target application; the edge device matches the target hard disk based on the preset matching rules. If the match is successful, the multimodal information is sent to the target hard disk for the target hard disk to store text information, voice information, and image information.

[0008] The present invention provides an intelligent conference recording and recording method that systematically addresses industry pain points, such as signal access restrictions and recording permission management. It supports access to mainstream video conferencing software, generates replay files driven by a timeline, and efficiently distributes them across multiple platforms. It employs an edge-cloud collaborative architecture, with edge preprocessing ensuring real-time performance, cloud-based deep analysis enhancing intelligence, and encryption strategies binding to device SNs ensuring data security, providing a complete technical system for intelligent conference scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 A flowchart of an intelligent conference record and recording method provided by an embodiment of the present invention when applied to an edge device;

[0010] Figure 2 A flowchart of an intelligent conference record and recording method provided by an embodiment of the present invention when applied to a cloud service device;

[0011] Figure 3 A module diagram of an intelligent conference record and recording system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0012] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention. Figure 1 To describe the system and method of intelligent conference recording and recording according to the exemplary embodiment of the present application. In one embodiment, the present application also proposes an intelligent conference recording and recording method. In the embodiment of the present application, an intelligent conference recording and recording method is applied to an edge device, such as Figure 1 As shown:

[0013] S101, obtaining audio information and video information.

[0014] In one embodiment, Bluetooth 5.0 and above are used, supporting the A2DP audio transmission protocol. 8 microphones can be paired simultaneously, and voice clarity is guaranteed through the SBC / AAC encoding format. The effective connection distance is up to 10 meters. 40kHz ultrasonic signals are used for device ranging and pairing. The microphone transmits an ultrasonic beacon. Yunbao receives the signal through the microphone array and calculates the phase difference to achieve centimeter-level positioning and automatic pairing, which is suitable for stable connections in complex electromagnetic environments. After the edge device turns on the pairing mode, it searches for WiFi / Bluetooth microphone devices synchronously through broadcast frames, and the ultrasonic microphone actively sends a ranging signal. The detected device is authenticated (such as a MAC address whitelist), and an encrypted transmission channel is established after verification (WiFi uses WPA2-PSK and Bluetooth uses AES-128 encryption).

[0015] For software like Zoom and Teams, access video stream data through official SDKs (such as the Zoom Web SDK and the Teams JavaScript SDK). Support for H.264 / H.265 encoding formats and resolutions up to 2160p / 60fps is supported. Leveraging edge device screen capture technology (such as the Windows Desktop Duplication API) allows real-time capture of PowerPoint sharing images, with adaptive dynamic refresh rates (15-60Hz). Using HDMI / USB interfaces, signals from external video conferencing devices (such as cameras and video conferencing endpoints) are converted into Yunbao audio and video data streams, supporting hot swapping and plug-and-play.

[0016] A 4-8 channel MEMS microphone array with a spacing of 10-15 cm uses beamforming technology to enhance the target sound source (speaker) signal and suppress ambient noise (such as keyboard sounds and air conditioning), improving the signal-to-noise ratio by 15-20 dB. The sampling rate is set to 44.1kHz / 16-bit, meeting CD-quality standards, and supports switching between mono and stereo modes. Both the participant's camera feed (1920×1080) and the PowerPoint presentation (1280×720) are captured simultaneously. Hardware encoding chips (such as Intel QuickSync) perform real-time compression to reduce bandwidth usage. A combination of spectral subtraction and Wiener filtering is used to remove both stationary noise (such as fan noise) and non-stationary noise (such as sudden coughing) from the voice signal. A 3D noise reduction algorithm is used for video to reduce noise in low-light environments. Voice signals are converted from PCM to FLAC / OPUS format (with a 2:1 compression ratio), and video is converted from YUV to H.264-encoded MP4 format for easy storage and processing. The audio and video clocks are synchronized through PTP (Precision Time Protocol), with the error controlled within ±5ms to avoid audio and video synchronization issues.

[0017] S102: Process the audio information and the video information in combination with the time axis to generate multimodal information, where the multimodal information includes text information, voice information, and image information.

[0018] In one implementation, edge devices use spectral subtraction and Wiener filtering to remove noise from collected voice signals and a 3D noise reduction algorithm to reduce noise in video. They then synchronize the audio and video clocks using the PTP protocol, with an error within ±5ms. For example, in a multinational conference, edge devices perform noise reduction on the audio and video of participants from different regions, then synchronize the processed audio and video based on a timeline to ensure synchronization. This processing generates multimodal information containing text, voice, and images. The edge device converts the audio into text and voice, and the video into images. For example, in a company's weekly meeting, the edge device converts the audio of participants' speeches into text using ASR technology, while retaining the voice information. The camera footage and PowerPoint presentation images are used as images, ultimately generating multimodal information containing text, voice, and images.

[0019] Before attending the meeting, enter a photo and voiceprint features to establish a database of facial feature vectors + voiceprint feature vectors + names. During the meeting, use a camera to capture the speaker's facial image in real time, extract facial features, and compare them with the database. Simultaneously, use a microphone to capture voice and extract voiceprint features, verifying identity (for example, identity is confirmed when the facial confidence level is >90% and the voiceprint match is >85%). If only the facial match is successful (e.g., the speaker is wearing a mask, causing voiceprint collection to fail), the speech record will be demarcated using the facial match result by default. If the voiceprint match is successful but there is no facial image (e.g., only the speaker spoke with voice), the voiceprint will be used to associate the identity. Edge devices can integrate lightweight facial recognition modules (e.g., real-time detection based on OpenCV) and voiceprint recognition SDKs (e.g., iFlytek Voiceprint Recognition). After completing feature extraction locally, the encrypted feature values ​​are sent to the cloud service device for database comparison. The cloud service device stores the participant information database. After the comparison results are returned to the edge device, they are associated with the timeline-driven speech-to-text results (such as "14:05:20 Zhang San: The focus of this meeting...").

[0020] When scheduling a meeting, the meeting creator can batch import participant photos, names, roles, and other information through the management background (such as importing from Excel templates). The system automatically generates a participant information database, which is suitable for internal fixed meetings (such as weekly meetings and monthly meetings). When corporate users create a meeting in the OA system, they simultaneously upload the list of participants and photos, and the system automatically associates them with the meeting ID. Participants are supported to independently enter photos and voiceprints through mobile APP, Web, and other entrances before the meeting begins (such as scanning the code to enter the meeting applet, taking a headshot and recording a 3-second voice). The system updates the information database in real time, which is suitable for temporary meetings or external guests. After entry, the system generates a time-sensitive QR code (integrated with the HMAC signature QR code in the document). Scan the code + face / voiceprint double verification when attending the meeting to improve identity accuracy.

[0021] In addition, a database of "facial feature vectors + names" can be established, and participant information can be imported in the following ways: Batch import (applicable to fixed meetings): The meeting creator uploads the photos, names, and role information of participants in batches using an Excel template through the management background (such as the OA system). The system extracts facial feature vectors through models such as FaceNet, associates them with the names, and stores them in the participant information database of the cloud service device to generate an identity information set corresponding to the meeting ID. Self-entry (applicable to temporary meetings): Before the meeting begins, participants scan the code through the mobile APP or Web to enter the mini program, take a front-facing photo, and enter their name. The system extracts facial features in real time and encrypts and stores them. At the same time, it generates a time-sensitive QR code with an HMAC signature (embedded with a facial feature hash value) for verification when attending the meeting. The edge device captures the video stream in real time through the UVC camera, uses OpenCV's MTCNN model to detect the face area, extracts the feature vector, and compares it with the locally cached feature library (or cloud service device interface). If the cosine similarity is greater than 0.85, the name is matched and the result is associated with the timeline-driven speech-to-text record (such as "14:05:20 Zhang San: The focus of this meeting...").

[0022] In addition, a database of facial feature vectors and names can be established. Participant information can be imported using the following methods: Batch import: When booking a meeting, enterprise users can upload attendee names and pre-recorded voice files (e.g., reading a specified text). The system extracts voiceprint feature vectors using MFCC or DeepSpeech models, binds them to the names, and stores them encrypted (AES-256) in a cloud service. Self-entry: Participants record a 3-second voice message on their mobile device (e.g., "I'm Zhang San, attending this meeting"). The system extracts voiceprint features and associates them with their names, updating the database in real time. The generated QR code contains a hash of the voiceprint features. The edge device frames the microphone input (50ms / frame), extracts voiceprint feature sequences using MFCC, and performs dynamic time warping (DTW) matching against a local or cloud-based voiceprint database. If the probability is greater than 0.7, the name is associated. The system then uses the document's 3D noise reduction algorithm to reduce ambient noise interference.

[0023] S103: Send the multimodal information to the cloud service device so that the cloud service device can generate target conference scene information.

[0024] In one implementation, after acquiring and processing audio and video information, the edge device generates multimodal information containing text, voice, and images, preparing to send it to the cloud service device. During a company's weekly meeting, the edge device connected to a microphone via WiFi and Bluetooth to capture audio, camera video, and PowerPoint presentations. After noise reduction, compression, and timeline synchronization, the device generated multimodal information containing the text of the meeting speech, voice clips, and images of the meeting screen.

[0025] Edge devices establish encrypted connections with cloud service devices using the TLS 1.3 secure communication protocol, listening on designated ports (such as 443) to create a secure channel for data transmission. For example, during a multinational company's monthly board meeting, edge devices establish TLS connections with cloud service devices at various branch offices to securely transmit multilingual meeting minutes. The edge device encodes the generated multimodal information (e.g., a 150MB meeting record containing encrypted audio and video clips and initial scene mode information) in chunks (1MB per chunk) and sends it to the cloud service device via the established secure channel. For example, after a company's weekly meeting, the edge device sends the chunked meeting record data to the cloud service device. The edge device calculates a SHA-256 hash value for each chunk and includes a checksum for the hash. If the cloud service device detects that the hash value of a data chunk (e.g., number 0x3A5F) is inconsistent with the locally calculated value, it receives a retransmission request and resends the chunk. If data loss or corruption occurs during transmission, the edge device performs retransmission and other actions based on the cloud service device's request to ensure data integrity. When the cloud service device finds that the 23rd block of data is missing when receiving data, the edge device receives a retransmission request and resends the missing data block to ensure that the cloud service device obtains complete multimodal information.

[0026] S104, performing matching processing with the target hard disk based on a preset matching rule, and if the match is successful, sending the multimodal information to the target hard disk for the target hard disk to store the text information, voice information and image information.

[0027] In one embodiment, when the target hard disk leaves the factory, a dedicated encryption tool is used to perform AES-256 encryption binding between the edge device serial number (such as "YUNBAO-20250622-001") and the hard disk unique identifier (such as "YUNBAO-HD-001") to generate an unalterable pairing record and store it in the hard disk firmware. The binding process uses hardware-level encryption (such as TPM chip) to ensure that the physical mapping relationship between the serial number and the hard disk cannot be forged or tampered with. When the edge device is connected to the target hard disk, the system automatically triggers firmware-level verification: the edge device reads the binding serial number stored on the hard disk; performs a binary comparison between its own serial number and the read serial number (the error must be ≤0 bit); if the match is successful (similarity 100%), an encrypted transmission channel is established; if it fails, the hard disk refuses to write data and triggers an alarm.

[0028] Because serial number pairing is completed at the factory, when the edge device sends audio and video data to the cloud service device, there is no need to perform the matching process again. Instead, it is directly stored on the target hard drive according to the following steps: Encryption Channel Establishment: Dynamic session keys are generated using the AES-256 algorithm. Multimodal information (such as a 150MB meeting record) is encrypted in blocks (each block is 1MB), ensuring full encryption of the transmission link (TLS1.3 protocol). Data Writing: Encrypted data is written directly to the designated hard drive partition (such as "Meeting Records_Encryption Area") via the M.2 interface. The write speed must meet 100MB / s or above (compatible with high-speed solid-state drives).

[0029] Based on the preset matching rules for the encrypted binding of the serial number information of the edge device and the target hard disk, the target hard disk is verified for serial number pairing validity. If the match is successful, the multimodal information is sent to the target hard disk through the AES256 encryption channel, and an encrypted transmission log is generated synchronously. At the same time, the interface information for rebinding the serial number information of the new edge device to the hard disk is reserved when the device is damaged. Specifically, the edge device verifies the serial number pairing of the target hard disk based on the preset rules for the encrypted binding of its own serial number information and the target hard disk. The serial number of the edge device is "YUNBAO-20250622-001", and the target hard disk has been encrypted and bound to this serial number when it leaves the factory. When the edge device is connected to the target hard disk, the system automatically reads the binding serial number stored in the hard disk and compares it with its own serial number to verify the validity of the pairing.

[0030] If the serial numbers match, the edge device sends the multimodal information to the target drive via an AES256 encrypted channel. After a company's weekly meeting, the edge device generates multimodal information containing text, audio, and images (e.g., a 150MB meeting record). This data is encrypted using the AES256 algorithm and then transmitted to the target drive via a dedicated encrypted channel, ensuring that the data is protected from theft or tampering during transmission. An encrypted transmission log is also generated simultaneously, recording key information during the transmission process. This log includes information such as the transmission time (June 22, 2025, 14:30:00), data size (150MB), encryption algorithm (AES256), target drive serial number (YUNBAO-HD-001), and transmission status (success / failure), facilitating subsequent traceability and auditing.

[0031] The interface information for rebinding the serial number of the new edge device to the hard drive when the reserved device is damaged. When the edge device is damaged due to a malfunction, technicians can use the reserved interface to rebind the serial number of the new edge device (such as "YUNBAO-20250623-001") to the target hard drive, enabling data migration without replacing the hard drive, thereby improving device replacement efficiency. The edge device achieves secure data interaction with the target hard drive through the standardized process of "serial number verification-encrypted transmission-log generation-interface reservation." This process not only ensures the storage security of multimodal information, but also improves the flexibility of device operation and maintenance through the reserved interface mechanism. It is suitable for local encrypted storage scenarios of enterprise meeting records.

[0032] like Figure 2 As shown, an intelligent conference record and recording method is applied to a cloud service device, including:

[0033] S201, receiving multimodal information sent by an edge device, where the multimodal information includes text information, voice information, and image information.

[0034] In one implementation, a cloud service device establishes an encrypted connection with an edge device using a secure communication protocol (such as TLS 1.3), listens on a designated port to receive multimodal information, and uses block encoding to ensure stable transmission of large files. After a company's weekly meeting, the edge device sends a 150MB meeting transcript (including encrypted audio and video clips and initial scene mode information) to the cloud service device. The cloud service device receives the data through a load balancing node and verifies the TCP sequence number to ensure no data is lost (for example, detecting a missing block 23 triggers a retransmission).

[0035] The cloud service calculates the SHA-256 hash value of the received data packet, compares it with the hash checksum provided by the edge device, and verifies that the data format complies with the protocol standard. If the hash value of the received encrypted financial data fragment is inconsistent with the local calculation result (for example, a bit flip caused by network fluctuations), the cloud service device marks the fragment as "corrupted" and requests the edge device to retransmit the data block numbered 0x3A5F.

[0036] The target format meeting information is parsed, separating audio, text, video, and encrypted metadata. The initial scenario information is parsed to extract the meeting outline, PowerPoint presentation, mind map structure, and timeline markers. These are stored by modality in a distributed file system and linked to the index database. The cloud service receives the initial scenario information for the "quarterly strategic meeting," parses the mind map structure (including the "market analysis" branch), stores the text content in a designated path, and stores the PowerPoint and mind map JSON structures in the metadata directory. An index is then created in the index library linking "meeting topic, time, and content type."

[0037] The cloud service device extracts key metadata (such as device SN and timestamp) from encrypted meeting information, verifies the device SN's legitimacy, and decrypts encrypted fragments in layers based on participant role permissions. The cloud service device receives encrypted financial meeting minutes from an edge device, extracts the device SN as YUNBAO-20250622-001, and verifies its presence on the whitelist. Because the administrator has "Financial Data Access" permissions, they can decrypt the CFO's AES-256-encrypted speech clips. Regular employees can only view the AES-128-encrypted meeting outline. The cloud service device securely receives and pre-processes multimodal information from edge devices through a standardized process: link establishment → verification and parsing → categorized storage → permission verification. For example, during a multinational company's board meeting, the cloud service device receives an 800MB data packet containing multilingual meeting minutes. After SHA-256 verification, modality-separated storage, and permission-decrypted data, it is synchronized to the edge device's local cache, providing standardized input data for subsequent multimodal prioritization and cross-meeting retrieval. This process ensures data integrity, security, and traceability.

[0038] S202. Process the multimodal information, combine context-aware error correction, content segmentation, dynamic segmentation strategy and preset storage strategy, and generate the target format meeting information and initial meeting scene mode information sent by the meeting record information receiving edge device.

[0039] In one implementation, the cloud service device establishes an encrypted connection with an edge device (Yunbao) via a secure communication protocol (such as TLS 1.3). It listens on a designated port (such as 443) to receive meeting information in the target format and initial meeting scene mode information. Data transmission uses block encoding (1MB per block) to ensure stable transmission of large files. After a company's weekly meeting, the edge device sends the meeting minutes (150MB, including encrypted audio / video clips and initial scene mode information) to the cloud service device. The cloud service device receives the data through a load balancing node and verifies the TCP sequence number to ensure data is not lost (for example, if the 23rd block is missing, a retransmission request is triggered).

[0040] The cloud service device calculates the SHA-256 hash value of the received data packet and compares it with the hash checksum attached to the edge device (incompleteness is determined if the error is ≤ 1 bit). The cloud service device also verifies that the data format complies with the conference information protocol standard (for example, JSON schema verifies field integrity, and XML format verifies tag closure). If the hash value of the encrypted financial data fragment in the received target format conference information is inconsistent with the locally calculated result (for example, a bit flipped during transmission from the edge device due to network fluctuations), the cloud service device marks the fragment as "corrupted" and requests the edge device to retransmit the data block (number 0x3A5F).

[0041] The target meeting format is parsed, separating audio (.wav), text (.txt), video (.mp4), and encryption metadata (such as role encryption level). The initial scenario information is parsed to extract the meeting outline, PowerPoint mind map structure, and timeline markers (such as "00:15-00:30 Technical Solution Discussion"). The information is stored in the distributed file system (HDFS) by modality, with audio stored in the voice_pool directory and text in the text_pool directory. An associated index database (MongoDB) records file metadata (such as creation time and encryption type). The cloud service device receives the initial scenario information for the "quarterly strategy meeting," parses the mind map structure (including the "market analysis" and "goal planning" branches), stores the text content in the / conference / 20250622 / strategy / text path, and stores the PowerPoint and mind map JSON structures in / metadata / mindmap.json. An associated index is created in the index library for "meeting topic-time-content type."

[0042] Extract the key metadata for the encrypted meeting information (such as the device SN and timestamp in the combined key) and verify the legitimacy of the device SN (by checking whether the SN is on the whitelist through the enterprise device management system). Decrypt the encrypted fragments in layers, first using the device SN to decrypt the base layer, then decrypting the corresponding content based on the participant's role permissions (such as cloud service device administrator permissions) (for example, only the CFO can decrypt the financial data fragment). The cloud service device receives the encrypted financial meeting minutes sent by the edge device and extracts the device SN as YUNBAO-20250622-001, verifying that it is on the enterprise whitelist. Because the cloud service device administrator has "Financial Data Access" permissions, it can decrypt the CFO's speech fragment (AES-256 encryption), while ordinary employee accounts can only view the decrypted meeting outline (AES-128 encryption).

[0043] The cloud service device securely receives and pre-processes meeting information on edge devices through a standardized process: "link establishment → verification and analysis → classified storage → permission verification." For example, during a multinational company's monthly board meeting, the cloud service device first establishes a TLS connection with edge devices at each branch office, receiving a data packet (approximately 800MB) containing multilingual meeting minutes (English and Chinese) and initial scenario mode information. After ensuring data integrity through SHA-256, it parses the data into voice clips (speeches by CEO, directors, and other roles), multilingual text transcripts (automatically translated summaries), and PowerPoint videos, storing them modally in distributed storage nodes across Europe and Asia Pacific. For encrypted director decision clips (encrypted using AES-256 + device SN), the cloud service device decrypts them after verifying administrator privileges through the company's AD domain. The processed meeting information is then synchronized to the edge device for local caching, providing standardized input data for subsequent multimodal prioritization and cross-meeting search.

[0044] S203: Encrypt the conference record information based on the participant role information to generate target format conference information.

[0045] In one implementation, the role information of participants is mapped to permissions and rules are normalized through a role resolution process. A role sensitivity assessment model and a dynamic encryption arbitration algorithm are introduced to achieve structured processing of encryption policies. A hybrid model (HybridModel) that integrates a multi-layer neural network with a rule engine is used, combining supervised learning (such as historical encryption data training) and unsupervised learning (such as cluster analysis of role sensitivity). The input layer of the model includes basic role features: participant job titles (such as "CFO" and "Technical Director"), departments (such as "Finance Department" and "R&D Department"), and job levels (such as L7 Director and L4 Engineer); meeting content features: keywords extracted from meeting records (such as "financial data" and "budget"), PPT titles, and speech duration.

[0046] The Word2Vec word embedding model converts role names and keywords into vector representations (128 dimensions). The semantic similarity between the role and the content is calculated (for example, the cosine similarity between "CFO" and "financial data" is greater than 0.8). The rule engine predefines a sensitivity mapping table (for example, "Finance Department + Budget Discussion → Sensitivity Level 5"). A two-layer fully connected neural network (with 64 and 32 hidden layer neurons, respectively) takes feature vectors as input and outputs a sensitivity level (1-5, with 5 being the highest). The rule engine output (for example, forcing "financial data" related content to level 5) is combined with the neural network predictions, and the final level is generated through a weighted summation (rule weight 0.6, neural network weight 0.4). This ultimately generates a sensitivity level for each role-content pair (for example, "CFO - Financial Data Discussion → Level 5" and "Director - General Speech → Level 3").

[0047] Semantic similarity threshold: 0.7 (exceeding this triggers a high-sensitivity flag); neural network learning rate: 0.001, iteration count: 5000, loss function: cross-entropy; rule engine priority: predefined sensitive terms (such as "password" and "key") take precedence over neural network predictions. An arbitration model based on a priority queue, combined with a greedy strategy and conflict resolution rules, enables dynamic scheduling of encryption policies. Priority is defined as follows: sensitivity level (Level 5 → AES-256, Level 3 → AES-128, Level 1 → No encryption); role authority weight (CEO = 0.9, CFO = 0.8, general attendee = 0.5); data timeliness (speech in the last 10 minutes → weight +0.3).

[0048] When data from multiple roles conflicts (e.g., the CEO and CFO speak simultaneously), the system prioritizes encryption based on "sensitivity level x role authority weight" (CFO - Level 5 x 0.8 = 4.0 > CEO - Level 3 x 0.9 = 2.7, giving priority to the CFO's content). Non-preemptive scheduling ensures that encrypted segments are not interrupted, while new segments enter a priority queue (with a queue length threshold of 10; if exceeded, lower-priority segments are discarded). The sensitivity model is updated every five minutes based on changes in meeting content (e.g., if the discussion topic changes from "project planning" to "financial budgeting," the CFO's content is automatically upgraded to a higher sensitivity level). A manual intervention interface is supported, allowing administrators to force encryption of a specific role's content to Level 5.

[0049] For example, consider a company's board meeting. The roles include the CEO (L9, decision-making level), the CFO (L8, finance department), and the directors (L7, multiple departments). The key words in the meeting content are "annual budget," "profit distribution" (from the CFO's speech), and "strategic planning" (from the CEO's speech). Word2Vec calculates the semantic similarity between "CFO" and "annual budget" as 0.85, triggering the rule engine's predefined "Finance Department + Budget → Level 5" rule. The neural network evaluates the CEO's "strategic planning" content as Level 3 (due to its low keyword sensitivity). The final output is: the CFO's content is Level 5, and the CEO's content is Level 3.

[0050] Encryption priority is calculated: CFO (5 × 0.8 = 4.0) > CEO (3 × 0.9 = 2.7). AES-256 encryption is scheduled for the CFO's speech segment (e.g., 00:20-00:30), while AES-128 encryption is used for the CEO's segment. Since there is no other high-priority content in the same time period, there is no queue conflict, and the encryption policies are executed in order. After each meeting, new role-sensitivity data (e.g., "Marketing Director - Customer List → Level 4") is automatically added to the training set, and the Word2Vec word embeddings and neural network parameters are updated. Model accuracy is regularly evaluated (weekly) using the F1 score (target ≥ 0.9). If it falls below the threshold, full data retraining is triggered. Model parameters are encrypted and stored (using the device SN as the encryption key) to prevent unauthorized tampering with the sensitivity mapping rules.

[0051] Integrate with the meeting permission protocol standard to build a three-dimensional permission matrix based on role, rank, and department. Generate a standardized encryption rule set through an encryption benchmark conversion model, and establish a dynamic update mechanism for encryption policies. Integrate with the meeting permission protocol standard to build a three-dimensional "role-rank-department" matrix (e.g., "Technology Department - Senior Manager - Decision-Making Level"). Through the encryption benchmark conversion model, map the matrix into standardized rules (e.g., "Senior Managers in the Technology Department can access the unencrypted version of department meeting minutes"), and establish a dynamic update mechanism (e.g., automatically refresh rules when personnel ranks change). A cloud service device builds a matrix for a project meeting, with the following details: Role dimension: Project Manager (Decision-Making Level), Development Engineer (Execution Level); Rank dimension: Director (L7), Manager (L6); Department dimension: R&D Department, Product Department. "L7 Directors in the R&D Department can decrypt all department meeting minutes, while Development Engineers can only access snippets of their own speeches."

[0052] Using the meeting cycle as a time window, the roles, job levels, departments, and historical encryption processing records of participants are integrated to construct a multi-dimensional encryption feature matrix, achieving full-dimensional correlation of encryption policies. Using the meeting cycle as a time window, the roles, job levels, departments, and historical encryption records of participants (such as the encryption methods used in a department's past meetings) are integrated to construct a multi-dimensional feature matrix (rows: role, columns: sensitivity level, depth: department), achieving full-dimensional correlation of encryption policies. In a previous meeting, the "Product Manager" encrypted the "Requirements Review" content at level 4. When this role discusses the requirements again in the current meeting, the cloud service device retrieves the historical policy from the matrix, automatically applies level 4 encryption (such as combined key + sharded storage), and marks it as "requires key encryption."

[0053] Missing encryption rules are supplemented through the sensitive prediction model, and key encryption policies are extracted by combining with the dynamic permission adjustment mechanism. These policies are then integrated with the three-dimensional permission matrix features to generate meeting information in the target format. Missing rules are supplemented with the sensitive prediction model (for example, when adding the "intern" role, its encryption permission is predicted to be the lowest level); combined with the dynamic permission adjustment mechanism (for example, temporarily granting visitors "read-only encryption" permissions), key policies are extracted and integrated with the three-dimensional matrix to generate encrypted meeting information in the target format. For an ad hoc meeting, the "external consultant" role was added, and the cloud service device's prediction model supplemented the rule: "External consultants can only access the summary portion of the meeting minutes, and the summary is encrypted using AES-128." In the target information finally generated by the integration, consultant-related content is marked as "[Encrypted Summary]" and is accompanied by a permission description document.

[0054] Cloud service devices implement dynamic encryption based on participant identity through a closed-loop process: "role resolution → matrix construction → feature association → policy completion." For example, during a quarterly financial report meeting, a CFO's discussion of financial data is automatically identified as highly sensitive content, triggering the highest encryption rule within the "role-level-department" matrix (e.g., AES-256 + device SN binding). This system also references encryption methods used by similar roles in previous meetings to ensure data security during storage (encryption chip sharding) and transmission (combined keys). Ultimately, meeting information is generated in a target format that can only be decrypted by authorized devices.

[0055] S204: Process the target format meeting information, combine the timeline-driven speech-to-text conversion and analysis processing to generate initial meeting scenario mode information including meeting outline, PPT, and mind map.

[0056] In one implementation, a timeline parsing process is used to map the target format meeting information to a timeline and normalize its rules. A meeting content evaluation model and a dynamic analysis and arbitration algorithm are introduced to achieve structured integration of meeting information processing. Through the timeline parsing process, the cloud service device maps the target format meeting information (encrypted meeting minutes) to the timeline, introducing a meeting content evaluation model (e.g., assessing importance based on speech duration) and a dynamic analysis and arbitration algorithm (resolving timing conflicts) to achieve structured integration of information. During a project meeting, the cloud service device sorted the CFO's financial report (00:30-00:45) and the technical director's solution presentation (00:15-00:30) by timeline, assessing that the financial report, due to its involvement with budget-sensitive content, had higher priority than the solution presentation. These data were then integrated into a structured data block consisting of "time + speaker + content type."

[0057] Aligned with the meeting review protocol standards, a three-dimensional processing matrix of time nodes, content types, and analysis dimensions was constructed. A standardized set of analysis rules was generated through a content benchmarking conversion model, and a dynamic update mechanism for meeting information processing was established. Aligned with the meeting review protocol standards, a three-dimensional matrix of "time nodes, content types, and analysis dimensions" was constructed (e.g., "00:15-00:30 - technical plan - feasibility analysis"). Standardized analysis rules were generated through the content benchmarking conversion model (e.g., "automatically generate content summaries every 15 minutes"), and a dynamic update mechanism was established (e.g., automatically expanding matrix dimensions when new content types appear). The cloud service device constructed a matrix for the quarterly summary meeting, specifically: time nodes: 00:00-00:15 (opening), 00:15-00:30 (performance report); content types: data reports, plan discussion; analysis dimensions: completion level, risk points. The generation rule is as follows: "After the performance report phase, automatically extract key indicators (e.g., revenue growth rate) from the data reports and generate a summary."

[0058] Using the meeting cycle as a time window, the system integrates time nodes, content types, analysis dimensions, and historical meeting processing records to construct a multidimensional meeting feature matrix, enabling full-dimensional correlation of meeting information processing. Using the meeting cycle as a time window, the system integrates time nodes, content types, analysis dimensions, and historical meeting records (e.g., analysis methods for similar past meetings) to construct a multidimensional feature matrix (rows: time, columns: content type, depth: historical similarity), enabling full-dimensional correlation. In historical meetings, "product iteration discussions" are often accompanied by "risk assessment" analysis. During the current meeting from 01:00 to 01:15 for product iteration discussions, the cloud service device retrieves historical policies from the matrix, automatically adds "risk assessment" to the analysis dimension, and associates historical risk cases (e.g., "2025 Q1 iteration delay risk").

[0059] The content prediction model completes missing analysis rules, combines the dynamic timeline adjustment mechanism to extract key strategies, and integrates them with the three-dimensional processing matrix features to generate initial meeting scenario information, including a meeting outline, PowerPoint presentation, and mind map. The content prediction model completes missing analysis rules (for example, when adding the "Agile Development" content type, the prediction requires the addition of the "Iteration Cycle" analysis dimension). Combined with the dynamic timeline adjustment mechanism (for example, automatically compressing analysis time for non-critical content after a meeting is delayed), key strategies are extracted and integrated with the three-dimensional matrix to generate the meeting outline, PowerPoint presentation, and mind map. A temporary discussion session on "AI Technology Implementation" was added, and the prediction model for cloud service devices completed the rule: "AI technology discussions must include the 'Technology Maturity' and 'Cost Budget' analysis dimensions." The resulting mind map breaks this session into branches: "Technical Principles → Maturity Assessment → Budget Planning," and links cost data from historical AI projects.

[0060] The cloud service device implements intelligent, time-series processing of meeting information through a closed-loop process: "timeline analysis → matrix construction → feature association → strategy generation." For example, in an annual strategic meeting, the cloud service device first integrates the CEO's opening speech (00:00-00:15) and the planning reports of department directors (00:15-01:30) according to the timeline. It then uses a three-dimensional matrix to label content types such as "strategic goals" and "resource allocation." By combining analysis methods for similar content from previous meetings (such as the resource allocation ratios from last year's strategic meeting), it completes missing analytical dimensions (such as the impact assessment of "market environment changes"). Ultimately, it generates a meeting outline, PPT, and mind map containing timeline summaries, content associations, and historical reference data, providing structured initial scenario model information for subsequent in-depth cloud-based analysis.

[0061] S205 , processing the target format conference information and the initial conference scene mode information based on the priority processing strategy of the multimodal information to generate conference information data after priority sorting and strategy processing.

[0062] In one embodiment, feature extraction and statistical analysis are performed on the target format meeting information and initial meeting scene mode information to generate information type feature information, sensitivity feature information, meeting role feature information, modal distribution feature information, historical priority processing feature information, and multimodal association feature information. The cloud service device extracts features from the received target format meeting information and initial scene mode information, including: information type features: identifying types such as text (meeting minutes), voice (speech clips), and video (PPT sharing); sensitivity features: assessing sensitivity levels through keyword matching (such as "financial data" and "password") and role permissions; and modal distribution features: calculating the data volume proportion of each modality (e.g., voice accounts for 60% and video accounts for 30%).

[0063] During a financial institution's annual budget meeting, the cloud service extracted meeting information in the target format, including the text of a "quarterly profit and loss statement" (sensitivity level 5) and the CFO's speech (modal distribution 25%). The mind map in the initial scenario model information contained a "risk assessment" node (the information type was an analytical report). Information type features, sensitivity features, meeting role features, modal distribution features, historical priority processing features, and multimodal correlation features were processed to generate information priority prediction, modal sensitivity assessment, information type correlation features, and multimodal fusion assessment. Information priority prediction combines sensitivity and role permissions to calculate priority (e.g., sensitivity level 5 + CEO role = 90% priority). Multimodal fusion assessment analyzes the temporal synchronization of voice, text, and video (e.g., a time difference of ≤1 second between voice and PowerPoint page turning indicates high fusion). When processing the budget meeting information above, the cloud service device calculated that the "quarterly income statement" text had a sensitivity level of 5 and was associated with the CFO role, with a priority prediction of 95%. The time difference between the text and the CFO's speech was detected to be 0.8 seconds, and the multimodal fusion degree was assessed as "high" (score 85 / 100).

[0064] Based on information priority prediction information, modal sensitivity assessment information, information type association feature information, and multimodal fusion assessment information, abnormal data in the target format meeting information and initial meeting scene mode information is marked and filtered to generate priority abnormal data screening results, including the type of abnormal information, modality, meeting role, the time of the abnormality, and the degree of the abnormality. Abnormal data is marked based on the assessment results, such as low-sensitivity information being incorrectly encrypted (type abnormality) and modal data being lost (such as video clip damage). If it is found that an ordinary speech (sensitivity level 2) is mistakenly encrypted with AES-256 (normally should be AES-128), the cloud service device will mark the abnormality type as "encryption level error", the time of occurrence is the 30th minute of the meeting, and the degree of abnormality is "medium" (affecting storage efficiency).

[0065] The results of the priority anomaly data screening are integrated and quantified to generate a priority impact factor. The priority impact factor represents the weight of the information type's impact on priority, the difficulty of modal processing, the effectiveness of historical priority strategies, and future trends in priority processing. This generates prioritized and strategically processed meeting information data. The anomaly data is integrated to generate a priority impact factor (e.g., an encryption level error has an impact factor of 0.6). Priority queues are generated based on "sensitivity x role weight x integration." After integrating all anomalies in the meeting, the resulting priority impact factor is 0.4 (low impact). The final ranking is: CFO financial data (95% priority) > CEO strategic speech (80%) > general participant discussions (50%). Cloud service devices reschedule resources based on this queue, prioritizing high-priority data.

[0066] Cloud service equipment achieves intelligent processing of multimodal meeting information through a closed loop of "feature extraction → evaluation and analysis → anomaly screening → priority sorting." For example, in a cross-border M&A meeting, the cloud service equipment first extracts features such as the legal agreement text (sensitivity level 5), the CEO's English speech (modal distribution 35%), and the M&A process flow chart video (information type: visual data). The agreement text is calculated to have a priority of 98% due to its inclusion of the keyword "intellectual property" and its association with the role of General Counsel, and its time synchronization with the video is excellent (90% fusion). If a video clip is detected to have lost 10 seconds due to network transmission (a serious anomaly), an impact factor of 0.7 is generated, triggering a retransmission mechanism. Finally, after prioritization, the encrypted storage and translation of the legal agreement are prioritized, ensuring the real-time and security of critical information and providing a well-organized data foundation for subsequent cross-meeting retrieval and multimodal fusion.

[0067] S206 , performing semantic analysis, feature fusion, and cross-modal association processing on the conference information data to generate target conference scene mode information.

[0068] In one implementation, a cloud service device performs semantic analysis on prioritized meeting information data (including voice, text, and video). This includes: speech semantic extraction: using ASR technology to convert speech into text and extract keywords (e.g., "Q2 revenue growth"); text semantic analysis: using the BERT model to analyze meeting minutes and identify entity relationships (e.g., "R&D department → 20% budget increase"); and video semantic understanding: extracting text from presentations using optical character recognition (OCR) and combining it with image recognition to analyze the meaning of charts (e.g., a bar chart representing each department's performance). At a tech company's product launch, a cloud service device analyzed the CEO's speech, "The new product AI chip will go into mass production in Q3," extracting the keywords "AI chip" and "Q3 mass production." It also analyzed chip parameter charts in the presentation video (OCR identifying "computing power 200TOPS") and established a semantic association: "AI chip → computing power 200TOPS → mass production in the third quarter of 2025."

[0069] Synchronize the speech text, PPT page turn times, and video frames to a unified timeline (with an error of ≤500ms). Convert speech keywords, text entities, and video OCR results into vectors of uniform dimension (e.g., 300 dimensions). Use an attention mechanism to calculate weights (40% for speech, 35% for text, and 25% for video) to generate a fused feature vector. During the meeting, the CTO delivered a speech presentation, "AI algorithm optimization has increased recognition rate to 95%" (voice). Simultaneously, the PPT page turned to the "Technical Solution" page (video). The text minutes recorded "Algorithm Optimization → Recognition Rate 95%." The cloud service device aligned the timestamps of the three (00:25:10) and generated a fused feature vector: [AI algorithm, optimization, recognition rate 95%, technical solution], assigning weights of 0.4 for speech, 0.35 for text, and 0.25 for video.

[0070] Association rule algorithms (such as Apriori) are used to discover intermodal associations (e.g., "Financial data PPT → CFO speech → Sensitive encryption"). Using knowledge graph technology, cross-modal entities (e.g., "Product manager," "Requirements document," and "Development progress") are constructed into an association network, supporting path queries (e.g., "Requirements change → Impacted module → Person in charge"). Cloud service device analysis of a marketing department meeting: the voice message "User feedback regarding lag" (00:10:00), the text minutes "Lagged issue → Technical department solution" (00:10:15), and the PPT annotation "Optimization plan → Launch next week" (00:10:30). After association processing, a graph is generated: User feedback → Lag issue → Technical department → Optimization plan → Launch time, forming a complete problem-solving chain. The fused semantic network is converted into JSON format (including timeline, entity relationships, and modal weights). Based on a library of historical meeting patterns (e.g., "Product review meeting" and "Financial budget meeting"), the current scenario type is matched to generate a labeled target scenario pattern (e.g., "Product review meeting_v2.0").

[0071] S207: confirming the target application based on the target conference scene mode information, and sending the target format instruction information to the target application for processing by the target application.

[0072] In one embodiment, feature extraction and analysis are performed on the target conference scene pattern information to generate scene semantic feature information, application matching feature information, platform protocol compatibility feature information, scene type distribution feature information, conference application ratio feature information, and device interface parameter feature information. The cloud service device extracts features from the target conference scene pattern information, including: scene semantic features (such as the semantic keywords of "technical solution review meeting"); application matching features (such as features compatible with conferencing software such as Teams and Zoom); and platform protocol compatibility features (such as HTTP and WebSocket protocol support).

[0073] The target scenario mode information is displayed as "Quarterly Financial Review Meeting." The cloud service device extracts the semantic keywords "financial data" and "budget report." The application matching feature points to "Enterprise WeChat Meeting" (because historical data shows that this software is frequently used for financial meetings). The platform protocol compatibility feature is "Support for HTTPS encrypted transmission." The scenario semantic feature information, application matching feature information, platform protocol compatibility feature information, scenario type distribution feature information, meeting application ratio feature information, and device interface parameter feature information are processed to generate scenario matching probability information, application compatibility assessment information, scenario type association feature information, and interface parameter adaptation assessment information. The extracted features are processed to calculate the matching probability between the scenario and the application (e.g., based on a neural network model trained on historical data). Application compatibility (e.g., whether an application supports the encrypted format of meeting records) and interface parameter adaptability (e.g., whether the video resolution matches the device output) are evaluated.

[0074] The cloud service device calculates a 92% match probability between "Quarterly Financial Review Meeting" and "Enterprise WeChat Meeting." Because the application supports AES-256 encryption (compatible with the encrypted format of meeting records) and its interface parameters are compatible with the device's 1080p video output (85% parameter match), it generates a "High Compatibility" evaluation result. Based on the scenario matching probability information, application compatibility evaluation information, scenario type-related feature information, and interface parameter adaptation evaluation information, the cloud service device marks and filters abnormal data in the target meeting scenario mode information, generating scenario abnormality data screening results that include the abnormal scenario type, application scenario, interface link, abnormality occurrence time, and abnormality severity. Based on the compatibility evaluation results, abnormal data in the target scenario is marked, such as unsupported encryption protocols and interface parameter conflicts, generating abnormality data screening results (including abnormality type and occurrence time). If the target scenario mode information requires "real-time subtitle generation," but the cloud service device detects that the target application (e.g., an older version of video software) does not support this feature, the cloud service device marks the abnormality type as "Function Missing," the occurrence time as "At the Start of the Meeting," and the abnormality severity as "Severe" (impacting meeting efficiency).

[0075] The results of the scenario anomaly data screening are integrated and quantified to generate a scenario execution impact factor. This factor represents the impact of the scenario type on execution, the difficulty of adapting the application scenario, the compatibility of the device interface, and the impact trend of future scenario execution. Based on the target meeting scenario mode information, the target application is identified and the target format command information is sent to the target application for processing. The results of the anomaly data screening are integrated and quantified to generate a scenario execution impact factor (e.g., the impact of a missing function on the meeting is weighted as 0.7). Based on the impact factor, the target application is identified (with the application with the lowest impact factor being prioritized), and the command information is sent. The cloud service device assessed the "Enterprise WeChat Meeting" scenario execution impact factor as 0.3 (low impact), while another alternative application had an impact factor of 0.8 due to interface compatibility issues. Ultimately, the application was selected for the "Enterprise WeChat Meeting" scenario and a target format command information containing the encrypted meeting transcript was sent, along with adaptation parameters (e.g., H.264 video encoding format, 44.1kHz audio sampling rate).

[0076] Cloud service devices achieve intelligent adaptation of conference scenarios and applications through a closed-loop process of "feature extraction → adaptation assessment → anomaly screening → quantitative decision-making." For example, in a "new product launch" scenario, the cloud service device first extracts scene semantic features ("product demonstration," "multi-platform live streaming") and application matching features (adapting to Douyin Live, WeChat Video Account, etc.), calculates the matching probability of each application (Douyin Live matching probability is 95%), and detects that Douyin Live supports the RTMP streaming protocol (compatible with the device interface), but finds that its encryption level is insufficient (anomaly type "security risk"), with a quantitative impact factor of 0.5 (medium impact). The cloud service device then activates a backup plan, selecting an enterprise live streaming platform with an 85% match and a high encryption level (impact factor 0.2). It finally confirms the platform and sends a command message to ensure the security and compatibility of the meeting records during transmission and storage.

[0077] In another implementation, the cloud service device creates an encrypted shared space based on the target meeting scenario mode information. This allows for access permissions (e.g., read-only, download, edit) to be configured by third-party roles (e.g., partner, customer), and associated with the meeting video encryption policy (e.g., AES-256 encryption + device SN binding). For example, after a technology company held a technical review meeting with an external supplier, the cloud service device created a shared space named "2025Q2 Chip Review," configured "read-only" permissions for the supplier representative (allowing them to view only the meeting video summary) and "download + annotate" permissions for the company's internal technical team. All shared content automatically inherits the role encryption rules used during the meeting recording (e.g., clips involving core technology require secondary authentication and decryption by the supplier's responsible person).

[0078] Connecting to third-party platforms' identity authentication protocols (such as OAuth2.0 and SAML) verifies the legitimacy of designated third parties. Establishing an encrypted transmission channel via TLS1.3 ensures that videos cannot be tampered with or stolen during sharing. When sharing conference videos to Tencent WeChat for Work, the cloud service device first verifies the WeChat for Work account's corporate credentials (such as domain name binding and administrator authorization) via OAuth2.0. The encrypted video stream (1080p resolution, 2Mbps bitrate) is then transmitted via TLS1.3. During transmission, each 10MB data block is accompanied by a SHA-256 hash checksum to ensure data integrity.

[0079] Meeting videos are stored in timeline segments (e.g., 15-minute segments), and an index of accessible segments is dynamically generated based on third-party permissions. Access can also be revoked in real time (e.g., if sensitive content is discovered after sharing, third-party access can be immediately terminated). In a cross-border M&A meeting, the cloud service device segmented the 8-hour meeting video into 32 segments, granting legal counsel access to the segment related to "M&A terms discussion" (04:30-06:00), while automatically hiding all other segments. If a segment is later discovered to contain undisclosed financial data, the cloud service device will update the permission policy in real time and revoke access to that segment.

[0080] Sharing operation logs are generated, recording information such as third-party access time, IP address, video clip access history, and download counts, complying with GDPR and SME Security 2.0 compliance requirements. A cloud service device records a bank customer's sharing of a video from the "Annual Risk Control Meeting": At 3:30 PM on June 26, 2025, a third party (XX Accounting Firm) accessed the video clip from 02:15 to 02:30 (containing a risk assessment report) through IP address 192.168.1.100 and downloaded it once. The log is automatically encrypted and stored (using the bank's proprietary key) for subsequent audit traceability. The video encoding format is automatically converted to a low-bitrate preview version (e.g., 720p, 500kbps) based on the technical standards of the third-party platform (e.g., TikTok Live's RTMP streaming and Zoom's MP4 format requirements) for quick preview. When the cloud service device shares the conference video to the Douyin enterprise account, it automatically converts the H.264-encoded MP4 file into the H.265-encoded TS format, adapts to Douyin's RTMP streaming protocol, and generates a 3-minute preview clip (including conference highlights) to facilitate rapid publishing on third-party platforms.

[0081] Cloud service devices enable secure sharing of conference videos with designated third parties through a comprehensive process encompassing "space creation - authentication - shard authorization - log auditing - and format adaptation." This process not only ensures data security during transmission and storage (e.g., AES-256 encryption and TLS tunneling), but also supports fine-grained, role-based permission control and real-time rights management. This makes it suitable for sharing conference materials between enterprises and third parties, including partners and customers, while also meeting compliance audit requirements.

[0082] S208, through semantic search technology, locates similar decision-making fragments in historical meetings and provides data support and processing capabilities for cross-meeting retrieval.

[0083] In one implementation, a semantic parsing process is used to perform semantic mapping and feature normalization on historical meeting decision fragments. A decision similarity assessment model and a cross-meeting association algorithm are introduced to achieve structured integration of decision information. Through the semantic parsing process, the cloud service device converts historical meeting decision fragments (such as the discussion record of "whether to launch an AI project") into semantic vectors. The BERT pre-trained model is used to extract keywords (such as "AI project," "budget," and "risk"), and Word2Vec is used to calculate word vector similarity (dimension 300). The decision fragment "Invest 5 million to develop an AI customer service system" from the Q4 2024 technical decision meeting was parsed, and the keyword vectors [AI customer service, 5 million, development, technical decision] were extracted. This was semantically mapped with the current meeting's decision on "AI chip R&D budget," resulting in a calculated similarity of 0.72.

[0084] The decision similarity assessment model uses a three-layer fully connected neural network (300-dimensional input layer, 128- and 64-dimensional hidden layers, and a 1-dimensional similarity score output layer). The loss function uses cosine similarity loss, a learning rate of 0.001, and 10,000 iterations. The model's input layer consists of word vectors and meeting metadata (time, role, and department). High-level semantic features are extracted using the ReLU activation function. The final output is a similarity score between 0 and 1 (with ≥0.6 indicating high similarity). The model assesses the similarity between the historical decision "AI customer service investment" and the current decision "AI chip R&D" to be 0.68. Because they share the semantic features of "AI technology" and "budget investment," this triggers a cross-meeting association.

[0085] Integrating with the conference semantic retrieval protocol standard, a three-dimensional index matrix of conference topics, decision types, and keywords was constructed. A standardized search index library was generated through a semantic benchmarking conversion model, and a search data update mechanism was established. Integrating with the semantic retrieval protocol, a three-dimensional index matrix (rows: conference topics, columns: decision types, depth: keywords) was constructed. For example, the topic dimension includes technical decision, financial budget, and product planning; the type dimension includes resource investment, risk assessment, and solution selection; and the keyword dimension includes AI, budget, and R&D cycle. An index was constructed for the "2024 AI Customer Service Decision Meeting": Topic → Technical Decision, Type → Resource Investment, Keywords → AI Customer Service, 5 million, and Development Cycle. Index items (technical decision, resource investment, AI Customer Service) were generated and stored in an HBase index table. Distributed search was achieved using a Lucene-based inverted index model combined with Elasticsearch, with real-time incremental updates (full refresh every 10 minutes). The tokenizer used was IKAnalyzer (supporting Chinese semantic token segmentation). The index had 6 shards (3 active and 3 standby), supporting petabyte-level data storage. Retrieval latency was in the millisecond range (99% of queries were ≤ 50ms).

[0086] Using the meeting lifecycle as a time window, we integrate meeting topics, decision types, keywords, and historical search records to construct a multidimensional semantic feature matrix, enabling full-dimensional correlation of decision information. Using the meeting lifecycle (e.g., from project initiation to completion) as a time window, we integrate topics (e.g., "annual budget"), types (e.g., "funding allocation"), keywords (e.g., "R&D funding"), and historical search records (e.g., if a keyword was searched 10 times) to construct an N×M×P-dimensional feature matrix (N = number of topics, M = number of types, P = number of keywords). We integrate the features of the "2024 Q4 Budget Meeting" and the "2025 Q1 Chip Meeting": both share the "Technical Budget" topic, the "Funding Allocation" type, and the shared keywords "R&D funding" and "Chips." The corresponding positions in the matrix are weighted upward (e.g., from 0.5 to 0.7).

[0087] The semantic completion model, based on a Transformer-based generative model (6-layer encoder, 6-layer decoder), takes as input a decision segment with missing features and outputs a completed semantic vector. The word embedding dimension is 768, the number of attention heads is 12, and the feedforward neural network dimension is 3072. Training data consists of 100,000 historical decision segments, with a BLEU-4 score of 0.85 or higher. A historical decision segment is missing the "budget amount" feature. Based on the context of "AI project" and "R&D cycle of 12 months," the model completes it with "budget of 8 million yuan." The similarity between this completed feature and the current meeting's "chip R&D budget of 10 million yuan" increases by 0.2.

[0088] A semantic completion model complements missing decision features, and a weighted distribution mechanism is used to extract key semantic information. This information is then integrated with the features of the three-dimensional index matrix to generate a standardized search result dataset, providing data support and processing capabilities for cross-conference searches. A weighted distribution mechanism (topic weight 0.4, type weight 0.3, keyword weight 0.3) extracts key semantics and integrates them with the three-dimensional index matrix to generate standardized search results (including the original decision fragment, similarity score, and associated tags). When searching for "AI chip budget," the weighted "chip" keyword was assigned a weight of 0.35 (higher than the default 0.3), matching the historical decision "AI customer service budget 5 million" with a similarity of 0.68. The result was labeled "Related to technical budget decision; reference to resource allocation ratio recommended."

[0089] The cloud service device achieves cross-meeting decision retrieval through a closed loop of "semantic analysis → index construction → feature fusion → result generation." For example, in the current "AI chip R&D decision meeting," the cloud service device: parses the decision fragment "Invest 10 million to develop AI chips" into a semantic vector, and uses a similarity model to calculate a similarity of 0.68 with the historical "AI customer service budget" decision; searches the three-dimensional index matrix to locate relevant records with the topic "technical decision," the type "resource investment," and the keyword "AI"; uses a semantic completion model to complete missing features of historical decisions (such as "technical risk assessment"), and constructs a multi-dimensional feature matrix;

[0090] Ultimately, standardized search results are generated, with recommendations based on historical budget allocation ratios (e.g., the adjustment basis from 5 million to 10 million) to provide data support for current decision-making. In this process, the similarity model's three-layer neural network ensures semantic matching accuracy, the Transformer completion model improves data integrity, and the distributed index library ensures search efficiency, forming a complete technical chain from data processing to application support.

[0091] like Figure 3 As shown, a system for intelligent conference recording and recording includes: obtaining audio information and video information by an edge device; processing the audio information and video information in combination with a timeline to generate multimodal information, the multimodal information including text information, voice information and image information; sending the multimodal information to a cloud service device for the cloud service device to generate target conference scene information;

[0092] The cloud service device receives multimodal information sent by the edge device, which includes text, voice, and image information. The multimodal information is processed and combined with context-aware error correction, content segmentation, dynamic segmentation strategy, and preset storage strategy to generate meeting record information. The meeting record information is encrypted based on the participant role information to generate target format meeting information. The target format meeting information is processed to generate initial meeting scene mode information including a meeting outline, PPT, and mind map. The target format meeting information and the initial meeting scene mode information are processed based on the priority processing strategy of the multimodal information to generate meeting information data after priority sorting and strategy processing. The meeting information data is subjected to semantic analysis, feature fusion, and cross-modal association processing to generate target meeting scene mode information. The target application is identified based on the target meeting scene mode information, and the target format instruction information is sent to the target application for processing by the target application. The edge device matches the target hard disk based on the preset matching rules. If the match is successful, the multimodal information is sent to the target hard disk for storage of the text, voice, and image information.

[0093] This solution's innovation lies in its multi-dimensional technology integration. Cloud service devices implement contextual error correction through lip-motion video and a thesaurus, combined with dynamic segmentation such as PPT page turning, and role-based encryption for security. A three-dimensional index matrix and a multi-dimensional feature matrix are constructed to enable semantic retrieval and multimodal fusion. The system addresses industry pain points such as signal access restrictions and recording permission management, supports access to mainstream video conferencing software, generates replay files driven by a timeline, and efficiently distributes across multiple platforms. It utilizes an edge-cloud collaborative architecture, with edge preprocessing ensuring real-time performance, deep cloud analysis enhancing intelligence, and encryption strategies tied to device SNs ensuring data security, providing a complete technical system for intelligent conferencing scenarios.

[0094] A computing device comprises a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute any one of the intelligent conference record and recording methods.

[0095] The methods and / or embodiments in the embodiments of the present application can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. When the computer program is executed by a processing unit, the above-mentioned functions defined in the method of the present application are performed.

[0096] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0097] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0098] It will be apparent to those skilled in the art that the present application is not limited to the details of the exemplary embodiments described above, and that the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the present application is defined by the appended claims rather than the foregoing description, and all variations that come within the meaning and range of equivalents of the claims are intended to be embraced herein.

Claims

1. An intelligent conference record and recording method, applied to cloud service equipment, characterized in that: include: Receive multimodal information sent by edge devices, including text information, voice information, and image information; Process multimodal information and generate meeting record information by combining context-aware error correction, content segmentation, dynamic segmentation strategy and preset storage strategy; Encrypt the meeting record information based on the participant role information to generate the target format meeting information; Process the target format meeting information to generate initial meeting scenario mode information including meeting outline, PPT, and mind map; The target format conference information and the initial conference scene mode information are processed based on the priority processing strategy of multimodal information to generate conference information data after priority sorting and strategy processing; Perform semantic analysis, feature fusion, and cross-modal association processing on conference information data to generate target conference scene mode information; Confirm the target application based on the target conference scene mode information, and send the target format instruction information to the target application for processing by the target application, including extracting and analyzing the feature of the target conference scene mode information to generate scene semantic feature information, application matching feature information, platform protocol compatibility feature information, scene type distribution feature information, conference application proportion feature information, and device interface parameter feature information; process the scene semantic feature information, application matching feature information, platform protocol compatibility feature information, scene type distribution feature information, conference application proportion feature information, and device interface parameter feature information to generate scene matching probability information, application compatibility evaluation information, scene type association feature information, and interface parameter adaptation evaluation information; based on the scene matching probability information, application compatibility evaluation information, scene type association feature information, and interface parameter adaptation evaluation information, mark and filter abnormal data in the target conference scene mode information to generate scene abnormality data filtering results, including the type of abnormal scene, application scenario, interface link, abnormality occurrence time, and abnormality degree; The results of scenario anomaly data screening are integrated and quantified to generate scenario execution impact factors. The scenario execution impact factors are used to characterize the impact weight of scenario type on execution, the difficulty of adapting the application scenario, the compatibility of device interfaces, and the impact trend of future scenario execution. Based on the target conference scenario mode information, the target application is identified and the target format instruction information is sent to the target application for processing. Through semantic search technology, similar decision-making fragments in historical meetings are located, providing data support and processing capabilities for cross-meeting retrieval.

2. The intelligent conference recording and recording method according to claim 1, characterized in that: Process multimodal information, combine context-aware error correction, content segmentation, dynamic segmentation strategy, and preset storage strategy to generate meeting record information, including: Collect and preprocess multimodal information to obtain original conference audio and video data; The original conference audio and video data is processed by combining lip movement video features with the conference theme vocabulary, performing context-aware error correction to generate processed conference audio and video data; Automatically segment and mark the processed conference audio and video data based on PPT page turning, time, and speaker information to complete content segmentation; Automatically segment the timeline based on speech pauses and speaker switching, mark silent segments based on discontinuous timeline marking technology, and convert non-speech information into text annotations; According to the preset storage strategy, the meeting records are stored in fragments in the device encryption chip to generate meeting record information.

3. The intelligent conference recording and recording method according to claim 1, characterized in that: Encrypt meeting records based on participant role information to generate target format meeting information, including: Through the role resolution process, permission mapping and rule normalization are performed on participant role information. A role sensitivity assessment model and dynamic encryption arbitration algorithm are introduced to achieve structured processing of encryption policies. Connect with the meeting permission protocol standard, build a three-dimensional permission matrix of role-level-department, generate a standardized encryption rule set through the encryption benchmark conversion model, and establish a dynamic update mechanism for encryption policies; Using the meeting cycle as a time window, we integrate the participant roles, job levels, departments, and historical encryption processing records to build a multi-dimensional encryption feature matrix, achieving full-dimensional correlation of encryption strategies. Missing encryption rules are supplemented by the sensitive prediction model, and key encryption strategies are extracted by combining with the dynamic permission adjustment mechanism. The strategies are then integrated with the three-dimensional permission matrix features to generate meeting information in the target format.

4. The intelligent conference recording and recording method according to claim 3, characterized in that: The target format meeting information is processed, combined with timeline-driven speech-to-text conversion and analysis processing to generate initial meeting scenario mode information including meeting outline, PPT, and mind map, including: Through the timeline parsing process, the target format meeting information is mapped to the time sequence and the rules are normalized. The meeting content evaluation model and dynamic analysis arbitration algorithm are introduced to achieve the structured integration of meeting information processing. Connect with the conference review protocol standards, build a three-dimensional processing matrix of time nodes, content types, and analysis dimensions, generate a standardized analysis rule set through the content benchmarking conversion model, and establish a dynamic update mechanism for conference information processing; Using the meeting cycle as a time window, integrating time nodes, content types, analysis dimensions, and historical meeting processing records, a multi-dimensional meeting feature matrix is ​​constructed to achieve full-dimensional correlation of meeting information processing; The missing analysis rules are completed through the content prediction model, and the key strategies are extracted by combining the dynamic adjustment mechanism of the timeline. They are then integrated with the three-dimensional processing matrix features to generate initial meeting scenario mode information including meeting outlines, PPTs, and mind maps.

5. The intelligent conference recording and recording method according to claim 1, characterized in that: The target format conference information and the initial conference scene mode information are processed based on the priority processing strategy of multimodal information to generate conference information data after priority sorting and strategy processing, including: Perform feature extraction and statistical analysis on the target format meeting information and initial meeting scene mode information to generate information type feature information, sensitivity feature information, meeting role feature information, modal distribution feature information, historical priority processing feature information, and multimodal association feature information; Process information type feature information, sensitivity feature information, conference role feature information, modal distribution feature information, historical priority processing feature information, and multimodal association feature information to generate information priority prediction information, modal sensitivity assessment information, information type association feature information, and multimodal fusion assessment information; Based on information priority prediction information, modal sensitivity assessment information, information type association feature information, and multimodal fusion assessment information, abnormal data in the target format meeting information and initial meeting scene mode information are marked and filtered, and priority abnormal data screening results are generated, including the type, modality, meeting role, abnormal occurrence time, and abnormality degree of the abnormal information; The priority anomaly data screening results are integrated and quantified to generate priority impact factors. The priority impact factors are used to characterize the influence weight of information type on priority, the processing difficulty of the modality, the effectiveness of historical priority strategies, and the impact trend of future priority processing, thereby generating meeting information data after priority sorting and strategy processing.

6. The intelligent conference recording and recording method according to claim 1, characterized in that: Using semantic search technology, we can locate similar decision-making fragments in historical meetings and provide data support and processing capabilities for cross-meeting retrieval, including: Through the semantic parsing process, historical meeting decision fragments are semantically mapped and feature normalized. A decision similarity evaluation model and a cross-meeting association algorithm are introduced to achieve structured integration of decision information. Connect with the conference semantic retrieval protocol standard, build a three-dimensional index matrix of conference topic-decision type-keyword, generate a standardized retrieval index library through the semantic benchmarking conversion model, and establish a retrieval data update mechanism; Using the meeting lifecycle as a time window, we integrate meeting topics, decision types, keywords, and historical search records to construct a multi-dimensional semantic feature matrix, achieving full-dimensional correlation of decision information. The semantic completion model is used to complete missing decision features, and the weight distribution mechanism is used to extract key semantic information. This information is then fused with the three-dimensional index matrix features to generate a retrieval result dataset in a standardized format, providing data support and processing capabilities for cross-conference retrieval.

7. An intelligent conference record and recording method, applied to edge devices, characterized in that: include: Obtain audio information and video information; Processing audio information and video information in combination with the time axis to generate multimodal information, which includes text information, voice information and image information; Sending the multimodal information to a cloud service device for the cloud service device to generate target conference scene information, wherein the cloud service device is configured to execute the method according to any one of claims 1 to 6; The multimodal information is matched with the target hard disk based on the preset matching rules. If the match is successful, the multimodal information is sent to the target hard disk for the target hard disk to store the text information, voice information and image information.

8. The intelligent conference recording and recording method according to claim 7, characterized in that: The target hard drive is matched based on the preset matching rules. If the match is successful, the multimodal information is sent to the target hard drive, including: Based on the preset matching rules of the serial number information of the edge device and the encrypted binding of the target hard disk, the serial number pairing validity of the target hard disk is verified. If the match is successful, the multimodal information is sent to the target hard disk through the AES256 encrypted channel, and an encrypted transmission log is generated simultaneously.

9. An intelligent conference record and recording system, characterized in that: include: The edge device obtains audio and video information; Processing the audio and video information in conjunction with the timeline to generate multimodal information, which includes text, voice, and image information; sending the multimodal information to the cloud service device so that the cloud service device can generate target conference scene information; The cloud service device receives multimodal information sent by the edge device, and the multimodal information includes text information, voice information, and image information; Process multimodal information and generate meeting record information by combining context-aware error correction, content segmentation, dynamic segmentation strategy and preset storage strategy; Encrypt the meeting record information based on the participant role information to generate the target format meeting information; Process the target format meeting information to generate initial meeting scenario mode information including meeting outline, PPT, and mind map; The target format conference information and the initial conference scene mode information are processed based on the priority processing strategy of multimodal information to generate conference information data after priority sorting and strategy processing; Perform semantic analysis, feature fusion, and cross-modal association processing on conference information data to generate target conference scene mode information; Confirm the target application based on the target conference scene mode information, and send the target format instruction information to the target application for processing by the target application, including extracting and analyzing the feature of the target conference scene mode information to generate scene semantic feature information, application matching feature information, platform protocol compatibility feature information, scene type distribution feature information, conference application proportion feature information, and device interface parameter feature information; process the scene semantic feature information, application matching feature information, platform protocol compatibility feature information, scene type distribution feature information, conference application proportion feature information, and device interface parameter feature information to generate scene matching probability information, application compatibility evaluation information, scene type association feature information, and interface parameter adaptation evaluation information; based on the scene matching probability information, application compatibility evaluation information, scene type association feature information, and interface parameter adaptation evaluation information, mark and filter abnormal data in the target conference scene mode information to generate scene abnormality data filtering results, including the type of abnormal scene, application scenario, interface link, abnormality occurrence time, and abnormality degree; The results of scenario anomaly data screening are integrated and quantified to generate scenario execution impact factors. The scenario execution impact factors are used to characterize the impact weight of scenario type on execution, the difficulty of adapting the application scenario, the compatibility of device interfaces, and the impact trend of future scenario execution. Based on the target conference scenario mode information, the target application is identified and the target format instruction information is sent to the target application for processing. The edge device matches the target hard disk based on preset matching rules. If the match is successful, the multimodal information is sent to the target hard disk for the target hard disk to store text information, voice information and image information.

Citation Information

Patent Citations

  • Encryption authentication method and device based on solid state disk, computer equipment and medium

    CN117235818A

  • AI multi-mode fusion interaction method, device, system and equipment

    CN120179079A