Conference live broadcast method and system

Through the collaboration of edge devices and cloud service devices, the problems of single control channel and stringent equipment requirements of video conferencing systems are solved, multimodal processing and security assurance are achieved, the flexibility of live broadcast and user experience are improved, and it is suitable for various scenarios such as medical care, education, and corporate business.

CN120602691BActive Publication Date: 2025-10-03XIAMEN RGBLINK SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511072068.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-10-03
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

Existing video conferencing systems have a single control channel and are unable to meet the diversified needs of the mobile Internet. They have stringent network and equipment requirements, and the semantic analysis and multimodal feature fusion are imperfect, resulting in poor live broadcast effects.

Method used

Through the collaborative work of edge devices and cloud service devices, dynamic matching of image quality between user roles and member information is achieved, and multimodal processing and security assurance are carried out, including receiving video signals, generating time-sensitive QR codes, processing meeting records and live broadcast information, and combining voice-to-text and cross-modal association to generate target meeting scene mode information.

Benefits of technology

It improves the flexibility, security and user experience of conference live broadcasts, and is suitable for various scenarios such as medical care, education, and corporate business. It achieves dynamic matching of image quality and security assurance based on user roles and member information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120602691B_ABST
    Figure CN120602691B_ABST
Patent Text Reader

Abstract

The present invention discloses a conference live broadcast method and system. The present invention comprises an edge device receiving a computer HDMI input signal and a UVC camera input signal of a video conference; sending live broadcast start information to a cloud service device, receiving a time-sensitive QR code and conference live broadcast information sent by the cloud service device that are signed with an HMAC and automatically expire within a preset time; obtaining target user conference attendance information based on the time-sensitive QR code; processing the target user conference attendance information to generate conference access rights, target user live broadcast image quality information and conference records with target user conference speech information; processing the conference live broadcast information based on the conference access rights and the target user live broadcast image quality information to generate conference live broadcast information; receiving target conference scene mode information sent by the cloud service device; confirming a target application based on the target conference scene mode information, and sending target format instruction information to the target application for processing by the target application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal data processing, and in particular to a conference live broadcast method and system. Background Art

[0002] Existing video conferencing systems still face numerous practical challenges. Some systems offer limited control channels for conferences, making them inadequate for the diverse needs of conference control brought about by the rise of mobile internet. In some scenarios, users simply need to participate in meetings as observers through live streaming, but most current video conferencing systems cannot meet such requirements. Furthermore, some solutions for implementing live video conferencing have stringent network and device requirements. For example, some high-definition video conferencing systems based on IMS networks have difficulty implementing live streaming over mobile networks, and selecting the live stream source from the conference client places excessive demands on the performance of the conference client. Other solutions use conference rooms as live streaming rooms, rather than live broadcasting the video conference. The live streamer must join the room via specific SIP messages, which limits the device type required to be a customized terminal supporting the SIP protocol.

[0003] When processing meeting information, existing technologies still need to be improved in areas such as semantic analysis, multimodal feature fusion, cross-modal association, and scene pattern generation. For example, it may be impossible to accurately extract semantic keywords and semantic paragraph information from meeting information, and it is difficult to efficiently achieve deep fusion and association of multimodal features such as voice, video, and text. The constructed cross-modal association network may contain logical irrationalities, resulting in the generated meeting scene pattern information failing to fully and accurately reflect the actual situation of the meeting, affecting the understanding, management, and application of meeting content. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0005] A conference live broadcast method, applied to an edge device, is characterized by comprising: receiving a computer HDMI input signal and a UVC camera input signal of a video conference; sending live broadcast start information to a cloud service device, receiving a time-sensitive QR code and conference live broadcast information sent by the cloud service device that have an HMAC signature and automatically expire within a preset time; obtaining target user conference attendance information based on the time-sensitive QR code; processing the target user conference attendance information to generate conference access rights, target user live broadcast image quality information, and a conference record with target user conference speech information; processing the conference live broadcast information based on the conference access rights and the target user live broadcast image quality information to generate conference live broadcast information; receiving target conference scene mode information sent by the cloud service device; confirming a target application based on the target conference scene mode information, and sending target format instruction information to the target application for processing by the target application.

[0006] A conference live broadcast method, applied to a cloud service device, is characterized by comprising: receiving signal input information and a request to start live broadcast sent by an edge device, generating a time-sensitive QR code with an HMAC signature that automatically expires within a preset time, and sending the code to the edge device; when illegal packet capture is detected, automatically cutting off the live stream and issuing an alarm, executing a live broadcast fuse mechanism, and physically separating the live stream from conference signaling; receiving conference record information sent by the edge device, encrypting the conference record information based on participant role information, and generating target format conference information; processing the target format conference information, combining timeline-driven speech-to-text conversion and analysis processing to generate initial conference scene mode information including a conference outline, PPT, and mind map; processing the target format conference information and the initial conference scene mode information based on a priority processing strategy for multimodal information to generate conference information data that has been prioritized and policy-processed; performing semantic analysis, feature fusion, and cross-modal association processing on the conference information data to generate target conference scene mode information, and returning the target conference scene mode information to the edge device for encrypted storage, while also performing local backup storage.

[0007] An intelligent conference live broadcast system includes: an edge device receives a computer HDMI input signal and a UVC camera input signal of a video conference, and sends a live broadcast start message to a cloud service device; a cloud service device receives the signal input information and a request to start the live broadcast sent by the edge device, generates a time-sensitive QR code with an HMAC signature and automatically expires within a preset time, and sends it to the edge device; when illegal packet capture is detected, the live broadcast stream is automatically cut off and an alarm is issued, and a live broadcast fuse mechanism is executed to physically separate the live broadcast stream from the conference signaling; the meeting record information sent by the edge device is received, the meeting record information is encrypted based on the role information of the participants, and a target format meeting information is generated; the target format meeting information is processed, and the initial meeting scene mode information including the meeting outline, PPT, and mind map is generated by combining the timeline-driven voice-to-text and analysis processing; the target format meeting information and the initial meeting scene mode information are prioritized based on the multimodal information processing strategy. Process the data to generate conference information data that has been prioritized and processed by policies; perform semantic analysis, feature fusion, and cross-modal association on the conference information data to generate target conference scene mode information, and return it to the edge device for encrypted storage, while performing local backup storage; the edge device receives the time-sensitive QR code and conference live broadcast information with an HMAC signature that automatically expires within a preset time sent by the cloud service device; obtains the target user's conference attendance information based on the time-sensitive QR code; processes the target user's conference attendance information to generate conference access rights, target user live broadcast image quality information, and conference records with target user's conference speech information; processes the conference live broadcast information based on the conference access rights and target user live broadcast image quality information to generate conference live broadcast information; receives the target conference scene mode information sent by the cloud service device; confirms the target application based on the target conference scene mode information, and sends the target format instruction information to the target application for processing by the target application.

[0008] The present invention provides a method for live conference streaming that addresses existing video conferencing systems' single control channel, stringent network and device requirements for live streaming, and semantic analysis in conference information processing. By collaborating with edge devices and cloud service devices, the system achieves dynamic image quality matching based on user roles and member information, multimodal processing of conference information, and security assurance. This improves the flexibility, security, and user experience of live conference streaming, making it suitable for a variety of scenarios requiring high-quality live streaming, such as healthcare, education, and enterprise commerce. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 A flowchart of a conference live broadcast method provided by an embodiment of the present invention when applied to an edge device;

[0010] Figure 2A flow chart of a conference live broadcast method provided by an embodiment of the present invention when applied to a cloud service device;

[0011] Figure 3 A module diagram of a system for a conference live broadcast method provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0012] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention and are not intended to limit the present invention. Figure 1 To describe the conference live broadcast method and system according to the exemplary embodiment of the present application. In one embodiment, the present application also proposes a conference live broadcast method.

[0013] In an embodiment of the present application, a conference live broadcast method is applied to an edge device, such as Figure 1 As shown:

[0014] S101, receiving the HDMI input signal of the computer for video conferencing and the input signal of the UVC camera.

[0015] In one implementation, the edge device receives computer video signals from software such as Tencent Video Conferencing via an HDMI interface, while simultaneously identifying and connecting to camera devices via the UVC (USB Video Class) standard interface. For example, when connecting to a Logitech C920 camera, the edge device automatically loads the UVC driver to complete device enumeration. The 1080p / 60fps video signal from the HDMI input is converted to H.264 encoding, while the YUV-format video data captured by the UVC camera undergoes resolution standardization. The input signal source resolution is converted based on the set live broadcast resolution, for example, downsampling a 4K camera input to 1080p to match the live broadcast requirements, or upscaling a 2K input source to 4K to match the live broadcast requirements. PTP (Precision Time Protocol) is used to calibrate the frame synchronization between the HDMI video signal and the UVC camera, eliminating audio and video asynchrony caused by device clock differences. For example, when a 20ms delay in the HDMI signal is detected, the camera data buffer is automatically adjusted to achieve synchronization.

[0016] Real-time monitoring of the HDMI signal's packet loss rate (e.g., triggering retransmission when the threshold is set to >5%) and the UVC camera's frame rate stability (e.g., enabling dynamic bitrate adjustment when the frame rate falls below 25 fps). De-noising filtering pre-processes the camera input, for example, using a 3D noise reduction algorithm to eliminate ambient noise. The processed HDMI and UVC signals are encapsulated into an RTMP stream and encrypted via TLS 1.3 for transmission to the cloud service device. For example, an AES-256 encryption header is embedded in the RTMP packet to prevent eavesdropping and tampering during transmission.

[0017] S102, sending live broadcast start information to the cloud service device, and receiving a time-sensitive QR code and conference live broadcast information sent by the cloud service device that has an HMAC signature and automatically expires within a preset time.

[0018] In one embodiment, the edge device encapsulates the basic information of the processed HDMI and UVC camera signals (such as resolution, encoding format, etc.), generates a request data packet containing the parameters required to start the live broadcast, and sends it to the cloud service device through a secure communication channel. For example, metadata such as the live broadcast title and expected start time are added to the data packet, which is encapsulated in JSON format and sent via the HTTPS protocol. The edge device receives the time-sensitive QR code with an HMAC signature returned by the cloud service device. The HMAC signature of the QR code is verified to ensure that it has not been tampered with, and its timeliness is checked (valid for 15 minutes). For example, the signature data in the QR code is hashed using the key shared with the cloud service device and compared with the received signature. If they are consistent and the current time is not more than 15 minutes beyond the generation time, the verification is successful.

[0019] Parse the conference live broadcast information from the response sent by the cloud service device, including the live stream address, playback key, live broadcast start time, etc. For example, extract the RTMP streaming address "rtmp: / / aabb.com / meeting / 12345" and playback key "abc123" from the JSON-formatted response data.

[0020] The verified time-sensitive QR code and conference live broadcast information are stored locally, and a corresponding expiration management mechanism is set. For example, a record is created in the edge device's database, associating the QR code ID with the live broadcast information, and a timer is set to automatically mark the QR code as expired 15 minutes after it is generated. The edge device generates a visual display interface for the QR code, allowing users to access and share it with non-meeting participants. It also provides an interface for users to query the live broadcast information status. For example, a QR code image is displayed on the edge device's management page, with the live broadcast start time and viewing link displayed below. Users can click to copy the link to share it.

[0021] S103, obtaining the target user's conference attendance information based on the timeliness QR code.

[0022] In one implementation, the edge device uses a built-in scanning module or an external scanning device to read the time-sensitive QR code scanned by the user, decode the QR code image, and extract the encrypted information contained therein. For example, a user uses a mobile phone to scan a QR code displayed by the edge device. The edge device's scanning interface receives the QR code image data and parses it using the ZBar decoding library to extract the URL link or parameter information.

[0023] Perform HMAC signature verification and timeliness checks on the decoded encrypted information to ensure it has not been tampered with and is valid for 15 minutes. For example, extract the encrypted string containing information such as the user ID and meeting ID from the QR code, perform HMAC signature verification using the key shared with the cloud service device, and compare the current time with the QR code generation time to confirm that the 15-minute validity period has not expired.

[0024] Parse the target user's meeting information from the verified QR code, including user ID, meeting ID, meeting time, and other content. For example, parse the parameters "user_id=12345&meeting_id=67890×tamp=1690000000" in the QR code to obtain the target user's meeting ID and meeting-related information. Based on the extracted user ID, obtain the target user's identity information and meeting attribute information, such as user role, department, and position, locally on the edge device or from the identity authentication system through an interface. For example, by calling the enterprise LDAP interface, based on user_id=12345, the query results in a user named "Zhang San", a role of "Intern", and a department of "Technical Department".

[0025] The target user's meeting attendance information is stored locally and associated with the live broadcast session for subsequent permission processing and live broadcast management. For example, a record is created in the edge device's database, linking information such as the user ID, meeting ID, and user role with the current live broadcast session ID. The attendance information is set to expire and automatically deleted after the live broadcast ends.

[0026] S104: Process the target user's conference participation information to generate conference access rights, target user's live broadcast quality information, and conference records containing the target user's conference speech information.

[0027] In one implementation, the target user's meeting attendance information is collected and verified to obtain the target user's identity and meeting attribute information. The edge device extracts the target user's basic meeting attendance information (such as user ID and meeting ID) from the time-sensitive QR code and verifies the legitimacy of the user's identity through the enterprise identity authentication system (such as LDAP or OAuth). The user ID is obtained from the QR code parameters "user_id = 1001 &meeting_id = 20240623" and the enterprise LDAP interface is called to query whether the user exists and is in a normal state. If the response is "user does not exist" or "account disabled," access is denied.

[0028] Dynamic permission isolation rules and audience classification strategies are combined to process the target user's identity and participant attributes, verify their eligibility, and generate a preliminary permission determination result. Combining dynamic permission isolation rules (such as prohibiting unauthorized users from speaking) with audience classification strategies (such as distinguishing between internal employees and external guests), user identities and participant attributes (such as department and meeting role) are verified to generate a preliminary permission determination result (such as "allow viewing" and "forbid interaction"). If the user role is "external guest," the permission rules automatically determine that they only have viewing permissions, without advanced permissions such as speaking and downloading, and generate a preliminary permission flag of "guest_view_only."

[0029] Based on the mapping rules between user role tags and live broadcast quality, the initial permission determination results are matched to image quality permissions and the quality level is set. User role tags represent the target user's job title and / or membership information. Based on the mapping rules between user role tags (such as "Intern" and "Expert") and live broadcast quality (such as Intern → 480p, Expert → 4K), the initial permission determination results are matched to the corresponding image quality level. If the user role tag is "Intern," the mapping table automatically sets the "480p" quality level, restricting the user from receiving HD video streams to avoid excessive resource consumption.

[0030] Permission boundary parameters are automatically generated based on conference security policies and device capability limits. Anomalous permission requests are flagged and prioritized based on a permission conflict detection mechanism. Permission boundary parameters (such as bandwidth limits and resolution limits) are automatically generated based on conference security policies (such as prohibiting unencrypted transmission) and edge device hardware capability limits (such as maximum decoding resolution). Anomalous permission requests (such as requesting 4K quality on a low-bandwidth network) are flagged and prioritized using a conflict detection mechanism. If the network bandwidth is detected as 1Mbps, the video quality permission limit is automatically adjusted to 720p. Requests for 4K quality are flagged as abnormal, forcibly downgraded, and prompted with the message "Insufficient network conditions, adjusted to optimal quality."

[0031] Based on pre-set permission generation rules, conference access permissions are associated with the target user's live broadcast quality information and encapsulated to generate permission configuration information with time-limited control. Edge devices semantically associate conference access permissions with live broadcast quality information based on pre-set permission generation rules (e.g., based on the RBAC role-based access control model). Access permissions cover basic functions (watching, downloading) and interactive functions (speaking, screen sharing). Quality information includes parameters such as resolution (480p / 720p / 1080p / 4K), bitrate (500kbps-8Mbps), and frame rate (25fps / 30fps).

[0032] If the user role is "Internal Expert," the rule automatically associates the "View + Interact" permission with the "4K / 8Mbps / 30fps" image quality configuration. A dual expiration control is implemented in the permission configuration. The live broadcast end time returned by the cloud service device (e.g., obtained through the metadata field in the RTMP protocol) is automatically populated into the expire_time field. If a user accesses via a QR code that expires in 15 minutes, the permission expiration period is automatically shortened to the remaining validity period of the QR code (e.g., if the user scans the code in the 10th minute, the permission expiration period is 5 minutes).

[0033] Specifically, for live medical surgery scenarios, since high-resolution images are required to display detailed information, the varying image quality requirements of different roles directly impact the efficiency of surgical teaching or collaboration. For example, if user A's role is labeled "Surgeon" and they have purchased a "Professional Membership" on a medical live streaming platform, according to the mapping rules, the surgeon needs to view tissue details in real time. Professional Members enjoy the highest image quality privileges, so they are matched with 4K / 8Mbps / 30fps (for example, showing the details of vascular sutures during laparoscopic surgery). For example, if user B's role is labeled "Intern" and they have not purchased a membership, since interns primarily observe and learn, their non-membership privileges limit high image quality, so they are matched with 720p / 1.5Mbps / 25fps (sufficient for viewing basic anatomical structures without consuming excessive bandwidth).

[0034] For live educational training scenarios, remote learning requires a balance between blackboard clarity and interactive flow. A membership system can differentiate between paid and free users regarding image quality. For example, consider user C, whose role is "Course Instructor" and who has purchased a Platinum membership. Since instructors need to present detailed presentations and handwritten blackboard notes, Platinum membership qualifies for the highest image quality, thus matching the video quality with 1080p / 4Mbps / 30fps (for example, zooming in on mathematical formula derivations). For user D, whose role is "Ordinary Student," who is a free registered user (without membership information), who only needs to watch basic courses, the free membership limits image quality, thus matching the video quality with 480p / 500kbps / 25fps (to meet text recognition requirements and reduce network load).

[0035] For corporate business meeting scenarios, high-level meetings need to ensure image stability and confidential information transmission, and membership levels can be associated with corporate customers' service priorities. When user E's role label is "Corporate Executive", the company he belongs to has purchased "Corporate Membership". Since executives participate in decision-making discussions, high-definition images are required to recognize facial expressions and body language. Corporate members enjoy exclusive bandwidth, so 4K / 6Mbps / 30fps image quality is matched (such as the simultaneous display of multiple venues in cross-border video conferences). When user F's role label is "Department Staff", the company has not purchased a membership (default basic permissions). Staff mainly receive information, and basic permissions limit non-essential image quality consumption, so 720p / 1Mbps / 25fps image quality is matched (sufficient for the speaker to have a clear image)

[0036] For the live broadcast scenario of academic seminars, the sharing of scientific research results requires accurate display of data charts and experimental processes, and membership services can distinguish between scientific research institutions and individual users. When the role label of user G is "Scientific Research Expert", the institution to which he belongs has purchased "Academic Professional Membership". Experts need to display paper charts and experimental videos. Professional members support high-resolution data display, so 1080p / 3Mbps / 30fps image quality is matched (such as real-time transmission of cell experiments under a microscope). When the role label of user H is "Student", the individual has not purchased a membership. Students mainly study by auditing, and non-members limit the image quality to ensure the core user experience, so 480p / 800kbps / 25fps image quality is matched (to meet the basic viewing needs of PPT text and the speaker's screen)

[0037] For the live broadcast scenario of financial investment, market analysis requires real-time refreshing of charts and data details, and membership level is directly related to service quality and information security. When user I's role label is "investment consultant", he has purchased "VIP membership". The consultant needs to analyze K-line charts and financial statements in real time, and VIP members enjoy low-latency and high-quality services, so 1080p / 5Mbps / 60fps image quality is matched (such as smooth refresh of stock market dynamic charts). When user J's role label is "ordinary investor", there is no membership information. Ordinary investors only need basic market display, and non-members limit the image quality to reduce server pressure, so 720p / 1.2Mbps / 30fps image quality is matched (to meet basic chart trend viewing). This application does not limit specific application scenarios. As long as there is a demand for live broadcast image quality, it belongs to the application scenario of this application.

[0038] Before attending the meeting, enter a photo and voiceprint features to establish a database of facial feature vectors + voiceprint feature vectors + names. During the meeting, use a camera to capture the speaker's facial image in real time, extract facial features, and compare them with the database. Simultaneously, use a microphone to capture voice and extract voiceprint features, verifying identity (for example, identity is confirmed when the facial confidence level is >90% and the voiceprint match is >85%). If only the facial match is successful (e.g., the speaker is wearing a mask, causing voiceprint collection to fail), the speech record will be demarcated using the facial match result by default. If the voiceprint match is successful but there is no facial image (e.g., only the speaker spoke with voice), the voiceprint will be used to associate the identity. Edge devices can integrate lightweight facial recognition modules (e.g., real-time detection based on OpenCV) and voiceprint recognition SDKs (e.g., iFlytek Voiceprint Recognition). After completing feature extraction locally, the encrypted feature values ​​are sent to the cloud service device for database comparison. The cloud service device stores the participant information database. After the comparison results are returned to the edge device, they are associated with the timeline-driven speech-to-text results (such as "14:05:20 Zhang San: The focus of this meeting...").

[0039] When scheduling a meeting, the meeting creator can batch import participant photos, names, roles, and other information through the management background (such as importing from Excel templates). The system automatically generates a participant information database, which is suitable for internal fixed meetings (such as weekly meetings and monthly meetings). When corporate users create a meeting in the OA system, they simultaneously upload the list of participants and photos, and the system automatically associates them with the meeting ID. Participants are supported to independently enter photos and voiceprints through mobile APP, Web, and other entrances before the meeting begins (such as scanning the code to enter the meeting applet, taking a headshot and recording a 3-second voice). The system updates the information database in real time, which is suitable for temporary meetings or external guests. After entry, the system generates a time-sensitive QR code (integrated with the HMAC signature QR code in the document). Scan the code + face / voiceprint double verification when attending the meeting to improve identity accuracy.

[0040] In addition, a database of "facial feature vectors + names" can be established, and participant information can be imported in the following ways: Batch import (applicable to fixed meetings): The meeting creator uploads the photos, names, and role information of participants in batches using an Excel template through the management background (such as the OA system). The system extracts facial feature vectors through models such as FaceNet, associates them with the names, and stores them in the participant information database of the cloud service device to generate an identity information set corresponding to the meeting ID. Self-entry (applicable to temporary meetings): Before the meeting begins, participants scan the code through the mobile APP or Web to enter the mini program, take a front-facing photo, and enter their name. The system extracts facial features in real time and encrypts and stores them. At the same time, it generates a time-sensitive QR code with an HMAC signature (embedded with a facial feature hash value) for verification when attending the meeting. The edge device captures the video stream in real time through the UVC camera, uses OpenCV's MTCNN model to detect the face area, extracts the feature vector, and compares it with the locally cached feature library (or cloud service device interface). If the cosine similarity is greater than 0.85, the name is matched and the result is associated with the timeline-driven speech-to-text record (such as "14:05:20 Zhang San: The focus of this meeting...").

[0041] In addition, a database of facial feature vectors and names can be established. Participant information can be imported using the following methods: Batch import: When booking a meeting, enterprise users can upload attendee names and pre-recorded voice files (e.g., reading a specified text). The system extracts voiceprint feature vectors using MFCC or DeepSpeech models, binds them to the names, and stores them encrypted (AES-256) in a cloud service. Self-entry: Participants record a 3-second voice message on their mobile device (e.g., "I'm Zhang San, attending this meeting"). The system extracts voiceprint features and associates them with their names, updating the database in real time. The generated QR code contains a hash of the voiceprint features. The edge device frames the microphone input (50ms / frame), extracts voiceprint feature sequences using MFCC, and performs dynamic time warping (DTW) matching against a local or cloud-based voiceprint database. If the probability is greater than 0.7, the name is associated. The system then uses the document's 3D noise reduction algorithm to reduce ambient noise interference.

[0042] Data packets are structured and encapsulated using a layered architecture. Specifically, the base layer contains the user unique identifier (UUID), the meeting ID (e.g., meeting_20240623), and the permission version number (V1.0). The function layer nests the access permission object (access_level) and the image quality configuration object (quality_config). The control layer includes the expiration parameter (expire_time), the signature algorithm (HS256), and the encryption key version (Key_V3).

[0043] The encryption is done using the JWT (JSON Web Token) specification and is divided into three parts:

[0044] Header: specifies the algorithm (for example, {"alg":"HS256","typ":"JWT"}).

[0045] Payload: stores the plain text of the permission configuration (e.g., {"user_id":"1001","access":"view interaction","quality":"1080p","exp":1700728800}).

[0046] Signature: Use the shared key between the edge device and the cloud service device (such as meeting_secret_2024) to perform HMAC-SHA256 signature on the first two parts to prevent tampering.

[0047] The web side transmits it through HTTPHeader (Authorization:Bearer${JWT}); the mobile side encapsulates it into a binary frame (frame header + JWT length + JWT content) through the Socket protocol.

[0048] S105: Process the conference live broadcast information based on the conference access permission and the target user's live broadcast quality information to generate the conference live broadcast information.

[0049] In one implementation, the conference live stream information is parsed and extracted based on the conference access permissions and the target user's live stream quality information to obtain the basic live stream content and permission control parameters. Based on the conference access permissions (e.g., view, interact) and the target user's image quality level, the edge device parses the basic content (e.g., video frames, audio streams) from the original live stream returned by the cloud service device and extracts the corresponding permission control parameters (e.g., a screen recording prohibition flag and image quality restriction flag).

[0050] The RTMP stream is parsed to extract H.264-encoded video frames and AAC audio streams. The permission control fields {"no_recording":true,"quality_limited":"720p"} are also extracted to confirm that the user is limited to 720p viewing and recording is prohibited. The dynamic permission isolation mechanism and live broadcast security policies are combined to process the basic live broadcast content and permission control parameters, perform access rights verification, and generate preliminary processed live broadcast information. Combining the dynamic permission isolation mechanism (e.g., prohibiting unauthorized users from sending data) with live broadcast security policies (e.g., encrypted transmission), permission verification is performed on the basic live broadcast content, filtering out content that exceeds the user's permissions (e.g., high-resolution video streams and interactive function interfaces). If the user's permission is "Intern-480p," the edge device automatically discards video streams of 1080p and above and blocks WebRTC interface calls to prevent the user from turning on the microphone to speak, generating preliminary live broadcast information containing only 480p resolution.

[0051] Based on the mapping rules between the target user's live video quality information and live stream encoding parameters, the quality parameters of the initially processed live video information are matched to complete the live stream encoding configuration. Based on the mapping rules between the target user's video quality information (such as resolution and bitrate) and live stream encoding parameters, the video encoding parameters are adjusted to ensure that the output stream complies with user permissions. If the user's video quality level is "Expert-4K," the edge device retains the encoding parameters of the original 4K video stream (resolution 3840×2160, bitrate 8Mbps). If the user's video quality level is "Trainee-480p," the video is downsampled to 854×480 and compressed to 500kbps, matching the mapping rule {480p:{width:854,height:480,bitrate:500k}}.

[0052] Automatically generate live streaming parameters based on conference security requirements and network transmission conditions. A parameter conflict detection mechanism flags and optimizes any anomalous transmission parameters. Transmission parameters (such as the RTMP streaming address, port, and buffer size) are automatically generated based on conference security requirements (such as TLS encryption) and network transmission conditions (such as bandwidth and latency). Parameter conflicts (such as high bitrate and low bandwidth incompatibility) are also detected. If the network detection indicates a 1.5Mbps bandwidth, the live streaming parameters are automatically adjusted to {protocol: RTMP / TLS, bitrate: 1000k, buffer_size: 500ms}. If the user requests 4K quality (requiring 8Mbps), this is flagged as an anomaly, forcibly downgraded to 720p, and the parameters are adjusted to {bitrate: 1500k}.

[0053] According to the preset live broadcast information generation rules, the conference access rights, target user live broadcast quality information and live broadcast content are associated and packaged to generate conference live broadcast information with permission control and quality level. According to the preset rules, the conference access rights, quality information and live broadcast content are associated and packaged to generate a live broadcast data package with time control. The FLV format live broadcast package is generated, and the package structure is as follows:

[0054] {"header":{"access_level":"view", / / Access permission: view only;

[0055] "quality":"720p", / / Image quality level;

[0056] "expire_time":"2024-06-2318:00:00", / / Expiration control;

[0057] "encryption": "AES-256" / / encryption method},

[0058] "video_stream":{...}, / / H.264 encoded video stream;

[0059] "audio_stream":{...} / / AAC encoded audio stream}.

[0060] When streaming to a CDN via RTMP, add permission parameters to the URL: rtmp: / / aabb.com / meeting?token=abc123&quality=720p&exp=1700728800 to ensure that users can only receive live streams with corresponding permissions.

[0061] Before encapsulating the live stream, a dynamic watermark is superimposed based on the user ID to prevent screen recording leaks. If user permissions are modified during the live stream (for example, upgrading from "Intern" to "Expert"), the edge device pushes the new permission configuration {"quality":"4K"} via WebSocket and re-encodes the video stream. For example, after detecting the permission change, it starts a 4K encoding thread and switches the streaming parameters.

[0062] S106: Receive target conference scene mode information sent by the cloud service device.

[0063] In one embodiment, the received target conference scene mode information is parsed to verify the integrity and validity of the information, such as verifying whether the information has been tampered with through HMAC signature. The scene mode information {"scene_mode" : "technical_discussion","version":"1.0", "signature":"a1b2c3..."} is received in JSON format, and the information content is calculated by HMAC-SHA256 using a preset key, compared with the signature field, and continued to process after verification. The scene mode parameters such as conference type, participant roles, functional requirements, etc. are extracted from the verified information. {"scene_mode":"technical_discussion","participant_roles":["host","expert","audience"],"required_features":["live_translation","Q&A"]} is extracted from the information to clarify that the conference scene is a technical discussion and requires live translation and question-and-answer functions.

[0064] S107: confirming the target application based on the target conference scene mode information, and sending the target format instruction information to the target application for processing by the target application.

[0065] In one embodiment, the meeting scene requirements are parsed and extracted based on the target meeting scene mode information to obtain scene mode parameters and application matching rules. The edge device parses the meeting scene requirements (such as meeting type and functional modules) from the target meeting scene mode information returned by the cloud service device and extracts application matching rules (such as function-application mapping relationships). If the scene mode information is {"scene_mode":"remote_training","required_features":["live_translation","screen_share","Q&A"]}, the edge device extracts the required parameters ["remote_training", "live_translation","screen_share"], and the matching rules are {"live_translation":["AI translation engine","bilingual subtitle generator"],"screen_share":["remote desktop tool","WebRTC plug-in"]}.

[0066] The scenario mode parameters and application matching rules are processed based on the conference function requirements and the application capability library. Application matching is verified to generate a preliminary list of matching target applications. The scenario parameters are matched and verified based on the conference function requirements and the local application capability library (which stores supported functions, interface protocols, etc.) to generate a preliminary list of eligible applications. If the "AI Translation Engine" capability library supports "live_translation" and is compatible with the version, and "Remote Desktop Tool" supports "screen_share", a candidate list is generated: "AI Translation Engine V2.1", "Remote Desktop Tool V3.5", "WebRTC Plugin V1.2"], and the capability matching level of each application is marked (for example, AI Translation Engine matching level is 95%).

[0067] Based on the mapping rules between the target conference scenario information and the API protocol, the interface protocol is adapted for the preliminarily matched target applications, completing the API configuration. Based on the mapping rules between the target conference scenario information and the API protocol, the interface protocol is adapted for the preliminarily matched applications (e.g., HTTP, WebSocket, RPC), completing the interface configuration. If the "AI Translation Engine" requires a WebSocket interface to transmit real-time audio streams, the edge device maps the audio parameters in the scenario information to {"audio_format":"PCM","sample_rate":16000}, configures the interface URL to wss: / / aabbcc.com / stream, and generates an authentication token (token=abc123) for permission verification.

[0068] Automatically generates command transmission parameters based on conference command transmission requirements and target application characteristics. A parameter compatibility detection mechanism flags abnormal transmission parameters and adapts them accordingly. Automatically generates transmission parameters (such as port number and timeout period) based on conference command transmission requirements (such as real-time performance and encryption level) and target application characteristics (such as parameter format and transmission protocol), and detects parameter conflicts. Generates transmission parameters such as {"protocol":"HTTPS","port":443,"timeout":3000,"encryption":"AES-256"}. If the application only supports HTTP / 1.1 and the scenario requires HTTPS, SSL tunnel adaptation is automatically enabled, adjusting the parameters to {"protocol":"HTTP","proxy":"https_proxy:8080"} and marking "encryption":"TLS1.2" for security.

[0069] According to the preset instruction generation rules, the target conference scene mode information is associated with the target format instruction information and encapsulated to generate an instruction data packet with scene adaptation and format specifications, and sent to the target application. According to the preset instruction generation rules, the target conference scene mode information is associated with the target format instruction information and encapsulated to generate an instruction data packet with scene adaptation, and sent to the target application. It is encapsulated into a JSON format data packet, as follows: {"scene_info":{"mode":"remote_training",

[0070] "timestamp":1690000000,"version":"1.0"},

[0071] "app_config":{"name":"AI Translation Engine V2.1","interface":"websocket",

[0072] "params":{"language":"zh-CN / en-US","output_format":"subtitle",

[0073] "delay_threshold": 500 / / delay threshold 500ms}},

[0074] "command":"start_translation",

[0075] "signature": "a1b2c3d4e5f6..." / / HMAC-SHA256 signature}.

[0076] The data is sent to wss: / / aabbcc.com / stream via WebSocket. The application parses the data and starts generating bilingual subtitles.

[0077] like Figure 2 As shown, a conference live broadcast method is applied to a cloud service device, comprising:

[0078] S201, receiving signal input information and a request to start live broadcasting sent by an edge device, generating a time-sensitive QR code with an HMAC signature that automatically expires within a preset time, and sending it to the edge device.

[0079] In one implementation, a cloud service device receives a live broadcast initiation request from an edge device via a secure communication channel, extracts signal input parameters (such as HDMI video resolution and UVC camera model) and basic conference information (such as the conference ID and live broadcast theme), and verifies the legitimacy of the request (such as the request format and permission authentication). Upon receiving a request from the edge device containing "Conference ID: 20240701-001, HDMI signal parameters: 1080P / 30fps, UVC camera model: Logitech C920, Live broadcast start time: 2024-07-01 14:00:00," the cloud service device verifies whether the edge device is on the authorized list and rejects the request if it is not. The QR code's core parameters, including the conference access path, user permission identifier, conference ID, and expiration timestamp, are assembled. The parameters are signed using the HMAC-SHA256 algorithm to generate an anti-counterfeiting signature (a signing key must be shared with the edge device). The parameters and signature are then encoded into a string recognizable by the QR code. Core parameters: meeting_id = 20240701-001 & access = view & expiration = 1698732000 & nonce = ts890 (where expiration is a timestamp 15 minutes in the future); use the key "cloud_live_secret" to generate a signature string sig = abc123def456; the final QR code content is: meeting_id = 20240701-001 & access = view & expiration = 1698732000 & nonce = ts890 & sig = abc123def456. The QR code is valid for 15 minutes, starting from the moment it is generated, and will automatically expire after the timeout. Based on the meeting type or edge device request parameters, the QR code is associated with corresponding live broadcast permissions (such as viewing quality and interactive function restrictions). If the edge device request contains "audience type = intern", the permission associated with the QR code is "480P image quality, no speaking"; if the generation time is 2024-07-01 14:00:00, it is valid until 2024-07-01 14:15:00, corresponding to the timestamp 1698732000.

[0080] The response data packet contains the QR code, live stream address (e.g., RTMP push address), and playback key. The data packet is encrypted (e.g., with AES-256) to ensure secure transmission. The response includes the following: a QR code image link or Base64-encoded data; live stream information (RTMP address: rtmp: / / aabb.com / 20240701-001, playback key: key_789); and expiration date (valid until 2024-07-01 14:15:00).

[0081] The response data packet is sent to the edge device via HTTPS, and the cloud service device records the operation log (such as the generation time, edge device IP address, and meeting ID) for subsequent auditing or troubleshooting. The log entry is "2024-07-01 14:00:05, generated a time-sensitive QR code for meeting 20240701-001 for edge device 192.168.1.5, valid for 15 minutes."

[0082] S202, when illegal packet capture is detected, the live stream is automatically cut off and an alarm is issued, and the live fuse mechanism is executed to physically separate the live stream and conference signaling.

[0083] In one implementation, when illegal packet capture is detected, network transmission data is monitored and analyzed in real time to capture packet capture behavior characteristics and abnormal transmission data. The cloud service device uses a network traffic monitoring module to capture live streaming data in real time, analyze data packet characteristics such as protocol type, transmission frequency, and data content, and compare them with a pre-defined illegal packet capture signature database to identify abnormal behavior (such as high-frequency data packets and unauthorized protocol access).

[0084] When an IP address is detected sending more than 50 unencrypted RTMP packets within 10 seconds, and the packet content contains abnormal instructions (such as "get_stream_key"), the system determines it as suspected illegal packet capture and records the captured IP address (192.168.1.101), time (2024-07-01 14:05:30), and packet characteristics (RTMP protocol, high-frequency requests).

[0085] Combined with security protection rules and live streaming circuit breaking policies, the system processes packet capture behavior characteristics and abnormal transmission data, conducts security threat assessments, and generates circuit breaking trigger instructions. Combining pre-set security protection rules (such as prohibiting unauthorized IP access) with live streaming circuit breaking policies (such as triggering circuit breaking due to high-frequency packet capture), the system assesses the threat level of packet capture behavior (low / medium / high) and generates a trigger instruction if the circuit breaking threshold is reached.

[0086] If the packet capture behavior is judged to be "high threat" (such as a data packet carrying malicious code), the cloud service device generates a circuit breaker instruction based on the policy, including "circuit breaker type = emergency disconnection, impact range = global live stream, trigger time = 2024-07-01 14:05:32".

[0087] Based on the physical separation rules for live streaming and conference signaling, the live streaming transmission link and the conference signaling transmission link are isolated to achieve transmission channel separation. Based on the physical separation rules for live streaming and conference signaling, network routing configuration or firewall policies are used to isolate the transmission channels for live streaming (RTMP) and conference signaling (WebRTC), ensuring that the attack only affects the live streaming and not the signaling system. The cloud service device modifies the load balancer configuration to direct the live streaming to a dedicated RTMP server cluster (IP range 10.0.1.0 / 24) and the conference signaling to a separate WebRTC server (IP range 10.0.2.0 / 24), severing the network interconnection path between the two and preventing packet capture from spreading to the signaling system.

[0088] Automatically generate alarms based on livestream security requirements and the system's alarm mechanism. Abnormal alarm levels are marked and sent accordingly based on the alarm priority detection mechanism. Alarms containing packet capture details are generated based on livestream security requirements (e.g., real-time alerts for high-level meetings) and the system's alarm mechanism. Alarms are marked with the threat level (e.g., red emergency alert) and sent to administrators via multiple channels (email, SMS, system notifications). The alarm message "[Urgent] Illegal packet capture detected (IP: 192.168.1.101, Threat Level: High)" is generated, with the alarm level marked "red." The alert is then prioritized and sent to the security operations team's SMS notification system.

[0089] According to the preset circuit breaker execution rules, illegal packet capture information is associated with circuit breaker operation instructions and encapsulated to generate an operation data packet with security marking and circuit breaker control, automatically disconnecting the live stream and issuing an alarm. According to the preset circuit breaker execution rules, illegal packet capture information (such as IP, time, and packet characteristics) and circuit breaker operation instructions (such as disconnecting the stream and recording a log) are encapsulated into an operation data packet and automatically sent to the live stream management module, which executes the disconnection operation and records the circuit breaker log.

[0090] The encapsulated data packet contains "capture IP = 192.168.1.101, capture time = 2024-07-01 14:05:30, operation instruction = cut off stream & ban IP, circuit breaker status = executed". After being sent to the RTMP server, the server immediately terminates the corresponding live stream (stream ID: meeting_20240701) and records "circuit breaker execution time: 2024-07-01 14:05:35, stream cut off" in the log.

[0091] S203: Receive the conference record information sent by the edge device, encrypt the conference record information based on the participant role information, and generate target format conference information.

[0092] In one implementation, the cloud service device receives meeting record information, including audio and video streams, text records, timeline markers, and other data, from the edge device via a secure channel (e.g., HTTPS / TLS). The cloud service device verifies the data's integrity (e.g., through MD5 checksum verification) and source legitimacy (e.g., through edge device identity authentication). Upon receiving a meeting record data packet from the edge device containing "Meeting ID: 20240701-001, Audio and Video Data Size: 500MB, Text Record: 10MB," the cloud service device verifies the sender's identity using a pre-set edge device certificate. If the certificate is expired, the cloud service device rejects the packet.

[0093] Extract participant role information (such as host, expert, and audience) from the meeting transcript. The encryption algorithm and key level (such as AES-256 or AES-128) are determined based on the role's security level (for example, the host is a highly sensitive role). The extracted participant role list includes "Host (Engineer Zhang), Expert (Professor Li), and Audience (five interns)." The cloud service device determines that the host's voice recordings must be encrypted with AES-256, and the audience's text recordings must be encrypted with AES-128.

[0094] Segment meeting transcripts by participant role (e.g., separate audio clips by speaker ID); apply corresponding encryption strategies to data from different roles (e.g., add an additional key layer to highly sensitive role data); and generate encryption metadata (e.g., encryption algorithm, key expiration date) associated with the encrypted data. Encrypt the audio clips of the host, Mr. Zhang, using AES-256 encryption with the key "host_key_20240701" and mark them as "highly sensitive." Encrypt the intern's text chat logs using AES-128 encryption with the key "guest_key_20240701" and mark them as "lowly sensitive."

[0095] Encapsulate the encrypted audio, video, and text data along with encrypted metadata (such as role tags, encryption time, and decryption permissions) into a unified format (such as JSON or binary), ensuring that the target format supports subsequent decryption and scene mode generation. The encapsulated target format meeting information includes: encrypted video stream (AES-256, host clips); encrypted text transcript (AES-128, audience chat); and metadata: {"role_encryption":{"host":"AES-256","guest":"AES-128"},"decrypt_permission":["ddeeff.com"]}.

[0096] S204: Process the target format meeting information, combine the timeline-driven speech-to-text conversion and analysis processing to generate initial meeting scenario mode information including meeting outline, PPT, and mind map.

[0097] In one implementation, timeline-driven speech-to-text conversion and analysis are combined to collect and convert the voice content in meeting minutes in real time, obtaining timeline-marked text. The cloud service device extracts the voice stream from the encrypted meeting minutes and, using a timeline-driven speech-to-text engine, converts the real-time voice content into a timestamped text record, forming a "timeline-speech-to-text" mapping. The edge device receives the host's voice stream (time period 14:00:00-14:05:00) and converts it into text: "14:00:15 Engineer Zhang: Today's meeting will mainly discuss technical solutions... 14:02:30 Engineer Zhang: The first module requires focused testing..." Each text segment is accompanied by a precise timestamp.

[0098] Combining meeting content analysis with keyword extraction strategies, the textual information marked in the timeline is processed and segmented to generate preliminary transcript data. Natural language processing techniques are used to perform semantic analysis on the timeline textual information, identifying paragraph boundaries and keywords, segmenting the continuous text into semantically independent paragraphs, and extracting core keywords (such as technical solution and test module). The textual information is segmented as follows: Paragraph 1 (2:00:15 PM - 2:02:00 PM): Keywords "technical solution, design framework"; Paragraph 2 (2:02:30 PM - 2:05:00 PM): Keywords "module testing, risk points."

[0099] Based on the timeline nodes, speech paragraph information, and meeting outline generation rules, the preliminary transcript data was structured to complete the meeting outline framework. Based on the timeline nodes and speech paragraph logic, and following the pre-set outline generation rules (such as the three-tiered structure of "topic-subtopic-key points"), the segmented text was organized into a hierarchical meeting outline. Specifically, the meeting topic was: Technical Solution Discussion, Subtopic 1: Design Framework, Key Point 1: Core Module Composition; Key Point 2: Interface Design Specifications. Subtopic 2: Test Plan, Key Point 1: Module Test Scope; Key Point 2: Risk Analysis.

[0100] Automatically generate PPT and mind map nodes based on the meeting theme and content logic. Use the node relevance detection mechanism to flag abnormal associations and make logical adjustments. Automatically generate PPT and mind map nodes based on the meeting theme and content logic. Use the node relevance algorithm to detect abnormal associations (such as irrelevant speeches that jump off topics), and adjust the node hierarchy and connection relationships. The main node "Technical Solution" is associated with the subnodes "Design Framework" and "Test Plan." The "Design Framework" subnode is associated with "Core Module" and "Interface Specification." If a "Module Test" node is detected as incorrectly associated with the "Design Framework," it will be automatically adjusted to the "Test Plan" and marked as "Abnormal Association, Corrected."

[0101] According to preset scenario generation rules, the timeline-driven speech-to-text information is associated and integrated with the meeting outline, PowerPoint presentation, and mind map, generating initial meeting scenario information containing the meeting outline, PowerPoint presentation, and mind map. According to preset rules, the timeline text information, meeting outline, PowerPoint presentation, and mind map are integrated in three dimensions to form initial meeting scenario information containing structured content and logical relationships. This integrated initial scenario information includes: a timeline index: "14:02:30" can be used to locate the outline's "Test Plan - Risk Analysis" node and the PPT and mind map's "Test Plan" nodes; a bidirectional mapping between the outline, PowerPoint presentation, and mind map: the outline's "Design Framework" node corresponds to the mind map's "Design Framework" node and its subnodes; and metadata: {"scene_type":"Technical Discussion","version":"1.0","timestamp":1698732000}.

[0102] S205 , based on the priority processing strategy of the multimodal information, the target format conference information and the initial conference scene mode information are processed to generate conference information data after priority sorting and strategy processing.

[0103] In one implementation, feature extraction and statistical analysis are performed on the target format meeting information and initial meeting scene mode information to generate modality type features, security sensitivity features, meeting role permissions features, multimodal data distribution features, historical priority processing rule features, and cross-modal correlation strength features. The cloud service device performs feature analysis on the encrypted target format meeting information and initial scene mode information, extracting six major feature categories from the multimodal data and performing statistical analysis. The meeting records identify the voice, video, and text modalities. For example, in a "technical discussion" meeting, voice accounts for 60%, video (PPT) accounts for 30%, and text chat accounts for 10%. Data sensitivity levels are assigned based on participant roles, such as "high sensitivity" for the host's speech and "low sensitivity" for intern chats. A mapping between roles and permissions is extracted, such as "experts" have 4K video quality permissions, while "interns" only have 480p. The time distribution of data is analyzed. For example, the first 30 minutes of a morning meeting focus on technical solutions, while the last 20 minutes involve sensitive data. Priority policies for historical meetings are retrieved, such as "encrypting highly sensitive voice data first." Calculate the correlation between audio and video. For example, if a certain audio segment corresponds to the explanation on page 3 of the PPT, the correlation strength is 0.85 (out of a maximum score of 1).

[0104] The system processes information on modality type, security sensitivity, meeting role permissions, multimodal data distribution, historical priority processing rules, and cross-modal correlation strength to generate modality priority predictions, sensitive data processing risk assessments, role permission-modality correlations, and multimodal data fusion assessments. Based on the extracted features, an algorithmic model generates four types of priority assessment information, providing a basis for anomalous data screening. Based on the meeting type (technical discussion), the voice modality is predicted to have the highest priority (weight 0.5), followed by video (0.3), and text (0.2). The audio clip in which the host mentions "core code" is assessed as "high risk" and requires prior encryption. The expert role's 4K video stream is associated with the "technical solution explanation" audio modality, generating a "high-quality transmission" policy. The synchronization rate between the audio and PowerPoint video reaches 90%, resulting in an "excellent" fusion assessment, necessitating prioritization for combined processing.

[0105] Based on modality priority prediction information, sensitive data processing risk assessment information, role permission and modality association characteristics, and multimodal data fusion assessment information, abnormal data in the target format meeting information and the initial meeting scenario mode information is marked and filtered, generating priority abnormal data screening results, including the abnormal information's modality type, sensitivity level, associated meeting role, abnormal occurrence time, and impact level. Based on this assessment information, abnormal data in the meeting information is marked, filtering out data segments that do not comply with the priority strategy, and recording the abnormal attributes. A request for a 4K quality video stream from an intern account (permission conflict) was detected, marked as "abnormal modality request," with a sensitivity level of "medium," an associated role of "intern," an occurrence time of 2:20:15 PM, and an impact level of "resource abuse." A segment of unencrypted host voice data was marked as "security violation," with a sensitivity level of "high," an associated role of "host," an occurrence time of 2:35:00 PM, and an impact level of "information leakage risk."

[0106] The results of the priority anomaly data screening are integrated and quantified to generate a priority impact factor. The priority impact factor represents the influence of modality type on priority, the complexity of sensitive data processing, the effectiveness of historical priority policy execution, and future priority adjustment trends. This factor then generates meeting information data after priority sorting and policy processing. The anomaly data screening results are quantified into a priority impact factor, taking into account modality weight, sensitivity complexity, and other factors, to ultimately generate ranked meeting information data. The priority impact factor is calculated as follows: for an intern requesting 4K video quality: Factor = Modality Weight (0.2) × Permission Conflict Coefficient (1.5) × Historical Policy Deviation (1.2) = 0.36; for an unencrypted host voice: Factor = Modality Weight (0.5) × Sensitivity Complexity (2.0) × Historical Policy Deviation (1.0) = 1.0. According to the factors sorted from high to low, unencrypted voice (factor 1.0) is processed first, followed by abnormal image quality requests (0.36), generating a processing queue: encrypting the host's unencrypted voice clip (14:35:00); rejecting the intern's 4K request and downgrading it to 480p (14:20:15).

[0107] S206, perform semantic analysis, feature fusion and cross-modal association processing on the conference information data to generate target conference scene mode information, and return it to the edge device for encrypted storage, while performing local backup storage.

[0108] In one implementation, semantic analysis and feature extraction are performed on meeting information data. Natural language processing techniques are then used to semantically parse the text within the meeting information data to obtain semantic keywords and semantic paragraph information. Cloud service devices utilize natural language processing techniques to semantically parse the text within meeting information data (such as timeline text records and chat logs), extract keywords and semantic paragraphs, and establish the semantic structure of the text. The text records of the "Technical Solution Discussion" meeting were analyzed, extracting the keywords "core module, interface design, testing risk," and dividing the text into semantic paragraphs: Paragraph 1: "Today's discussion will focus on the core module components of the technical solution..." (Keywords: technical solution, core module); Paragraph 2: "Interface design must meet high concurrency requirements..." (Keywords: interface design, high concurrency).

[0109] Combining multimodal feature fusion strategies, semantic keywords and paragraph information are processed with speech and video features to perform cross-modal feature association and generate multimodal fusion feature data. Semantic analysis results are cross-modally associated with speech features (such as voiceprints and intonation) and video features (such as PPT images and speaker movements). This multimodal feature fusion strategy generates fused data and establishes relationships between different modalities.

[0110] The following are the speech features: the voiceprint features when the host mentions "core module" are associated with the text keyword "core module"; the following are the video features: the image features when the PPT opens to the "Module Architecture" page are associated with the speech paragraph "Now explain the module composition." A multimodal feature group for "Core Module Explanation" is generated, including text keywords, speech voiceprint fragments, and video frame screenshots.

[0111] Based on cross-modal association rules and the logical relationships between conference scenarios, multimodal fusion feature data is structured and integrated to complete the construction of a cross-modal association network. Based on cross-modal association rules and conference logic, the fused feature data is structured and integrated into an association network, with nodes representing features and edges representing logical relationships between features (such as causality and parallelism). The main node "Technical Solution" is associated with the sub-nodes "Core Module" and "Interface Design." The "Core Module" node connects the voice feature sub-node (the host's explanation audio), the video feature sub-node (the module architecture PPT page), and the text feature sub-node (the core module keyword paragraph). The edges between nodes mark the synchronization relationship between "explanation-image-text", for example, the time synchronization deviation between the "Core Module" voice and the PPT image is less than 500ms.

[0112] Based on the conference scene model generation rules and application requirements, the cross-modal association network is verified and abnormal associations are adjusted based on the scene model integrity detection mechanism to generate the target conference scene model framework. Based on the conference scene model generation rules and integrity detection mechanism, the logical rationality of the association network is verified, abnormal associations (such as incorrect connections between unrelated modalities) are adjusted, and the target conference scene model framework is generated.

[0113] It was detected that the "Test Risk" text paragraph was incorrectly associated with the "Interface Design" video node. According to the conference logic (test risk belongs to the test plan), it was adjusted to the "Test Plan" node. Verification found that the voice and video synchronization deviation associated with the "Core Module" was >1000ms, which was marked as an abnormality, and the synchronized voice and video clips were re-matched.

[0114] According to the preset scene pattern generation rules, the semantic analysis results, multimodal fusion feature data and cross-modal association network are associated and integrated to generate the target meeting scene pattern information. According to the preset rules, the semantic analysis results, multimodal fusion data and the optimized association network are integrated to generate the target meeting scene pattern information containing structured content and cross-modal association.

[0115] The integrated target scenario model information includes the following content and semantic structure: Technical Solution → Core Module → Module Composition, Interface Design → High Concurrency Requirements, Test Plan → Risk Points. The cross-modal association information is as follows: the semantic node represents the core module composition, the associated audio segment is the host's presentation from 14:05:10 to 14:08:30, and the associated video frame is the module architecture diagram on page 3 of the PPT.

[0116] like Figure 3 As shown, a system for recording and recording intelligent conferences includes: an edge device receives a computer HDMI input signal and a UVC camera input signal of a video conference, and sends a live broadcast start message to a cloud service device;

[0117] The cloud service device receives the signal input information and the request to start live broadcast sent by the edge device, generates a time-sensitive QR code with an HMAC signature that automatically expires within a preset time, and sends it to the edge device; when illegal packet capture is detected, the live stream is automatically cut off and an alarm is issued, and the live broadcast fuse mechanism is executed to physically separate the live stream from the conference signaling; the meeting record information sent by the edge device is received, and the meeting record information is encrypted based on the role information of the participants to generate the target format meeting information; the target format meeting information is processed, and the initial meeting scene mode information including the meeting outline and PPT is generated by combining the timeline-driven voice-to-text and analysis processing; based on the priority processing strategy of multimodal information, the target format meeting information and the initial meeting scene mode information are processed to generate meeting information data after priority sorting and strategy processing; the meeting information data is subjected to semantic analysis, feature fusion and cross-modal association processing to generate the target meeting scene mode information, and returned to the edge device for encrypted storage, while also performing local backup storage;

[0118] The edge device receives the time-sensitive QR code and conference live broadcast information sent by the cloud service device, which are signed with an HMAC and automatically expire within a preset time; obtains the target user's conference attendance information based on the time-sensitive QR code; processes the target user's conference attendance information to generate conference access rights and target user's live broadcast image quality information; processes the conference live broadcast information based on the conference access rights and target user's live broadcast image quality information to generate conference live broadcast information; receives the target conference scene mode information sent by the cloud service device; confirms the target application based on the target conference scene mode information, and sends the target format instruction information to the target application for processing by the target application.

[0119] This application is dedicated to solving the problems of the existing video conferencing system with a single control channel, strict requirements on the network and equipment for live broadcast, and semantic analysis in conference information processing. On the edge device side, the computer HDMI input signal and UVC camera input signal are first received, and after processing such as encoding format conversion and frame synchronization calibration, the live broadcast start information is sent to the cloud service device to obtain a time-sensitive QR code and conference live broadcast information with an HMAC signature that automatically expires in 15 minutes. Then, the target user's conference participation information is obtained based on the QR code, and the conference access rights and target user live broadcast image quality information are generated in combination with dynamic permission isolation rules. The conference live broadcast information is then processed to generate live broadcast information with permission control and image quality level. Finally, the target application is confirmed and instructions are sent based on the target conference scene mode information sent by the cloud service device.

[0120] On the cloud service device side, it generates a time-sensitive QR code after receiving the signal input information and live broadcast request from the edge device. When illegal packet capture is detected, the live broadcast fuse mechanism is executed to physically separate the live broadcast stream from the conference signaling. It receives the conference record information and encrypts it based on the role of the participants. It generates the initial conference scene mode information by combining the timeline to drive speech-to-text, etc. It generates the target conference scene mode information through multimodal information priority processing and semantic analysis, and returns it to the edge device for encrypted storage and local backup. Through the collaboration of edge devices and cloud service devices, the system realizes dynamic matching of image quality based on user roles and member information, multimodal processing of conference information and security assurance, which improves the flexibility, security and user experience of conference live broadcasts. It is suitable for various scenarios with high live broadcast quality requirements, such as medical care, education, and corporate business.

[0121] A computing device comprises a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute any one of the conference live broadcast methods.

[0122] The methods and / or embodiments in the embodiments of the present application can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. When the computer program is executed by a processing unit, the above-mentioned functions defined in the method of the present application are performed.

[0123] It should be noted that the computer-readable medium described in this application may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, a computer-readable medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0124] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0125] It will be apparent to those skilled in the art that the present application is not limited to the details of the exemplary embodiments described above, and that the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the present application is defined by the appended claims rather than the foregoing description, and all variations that come within the meaning and range of equivalents of the claims are intended to be embraced herein.

Claims

1. A conference live broadcast method, applied to edge devices, characterized in that: include: Receive computer HDMI input signals and UVC camera input signals for video conferencing; Send live broadcast start information to the cloud service device, and receive the time-sensitive QR code and conference live broadcast information sent by the cloud service device with an HMAC signature and automatically expiring within a preset time; Obtain target user meeting attendance information based on time-sensitive QR codes; Process the target user's meeting attendance information to generate meeting access rights, target user live broadcast quality information, and meeting records containing the target user's meeting speech information, including collecting and verifying the target user's meeting attendance information to obtain the target user's identity and meeting attribute information; Combine dynamic permission isolation rules with audience classification strategies to process target user identities and participant attribute information, perform participant qualification verification, and generate preliminary permission determination results; Based on the mapping rules between user role tags and live broadcast quality, the preliminary permission determination results are matched with image quality permissions to complete the image quality level setting. Among them, the user role tag is used to represent the target user's job position and / or membership information of the target user; permission boundary parameters are automatically generated according to the conference security policy and device capability limitations, and abnormal permission requests are marked and prioritized based on the permission conflict detection mechanism; according to the preset permission generation rules, the conference access rights are associated and packaged with the target user's live broadcast quality information to generate permission configuration information with time control; the target user's facial information before participating in the meeting is obtained; the target user's facial information before participating in the meeting and the meeting speech information at the time of participating in the meeting are processed to generate a meeting record containing the target user's meeting speech information; The conference live broadcast information is processed based on the conference access rights and the live broadcast quality information of the target user to generate the conference live broadcast information; Receive target conference scene mode information sent by the cloud service device; The target application is confirmed based on the target conference scene mode information, and the target format instruction information is sent to the target application for processing by the target application.

2. The conference live broadcast method according to claim 1, characterized in that: The conference live broadcast information is processed based on the conference access rights and the target user's live broadcast quality information to generate conference live broadcast information, including: Parse and extract conference live broadcast information based on conference access rights and target user live broadcast quality information to obtain basic live broadcast content and permission control parameters; Combine the dynamic permission isolation mechanism with the live broadcast security policy to process the basic live broadcast content and permission control parameters, perform access permission verification, and generate preliminarily processed live broadcast information; Based on the mapping rules between the target user's live broadcast quality information and the live broadcast encoding parameter, the live broadcast information after preliminary processing is matched with the quality parameters to complete the live broadcast encoding configuration; Automatically generate live streaming transmission parameters based on conference security requirements and network transmission conditions, and mark abnormal transmission parameters based on parameter conflict detection mechanism and optimize and adjust them; According to the preset live broadcast information generation rules, the conference access rights, target user live broadcast quality information and live broadcast content are associated and packaged to generate conference live broadcast information with permission control and quality level.

3. The conference live broadcast method according to claim 1, characterized in that: The target application is determined based on the target conference scene mode information, and the target format instruction information is sent to the target application for processing by the target application, including: Analyze and extract meeting scene requirements based on target meeting scene mode information, obtain scene mode parameters and apply matching rules; Combine the conference function requirements with the application capability library to process the scene mode parameters and application matching rules, perform application matching verification, and generate a preliminary matching target application list; Based on the target conference scenario mode information and the application interface protocol mapping rules, the interface protocol is adapted for the preliminary matched target application list to complete the application interface configuration; Automatically generate command transmission parameters based on conference command transmission requirements and target application characteristics, and mark abnormal transmission parameters based on parameter compatibility detection mechanism and make adaptive adjustments; According to the preset instruction generation rules, the target conference scene mode information is associated with the target format instruction information and packaged to generate an instruction data packet with scene adaptation and format specifications, and sent to the target application.

4. A conference live broadcast method, applied to a cloud service device, characterized in that: include: Receive the signal input information and live broadcast start request sent by the edge device, generate a time-sensitive QR code with an HMAC signature and automatically expires within a preset time, and send it to the edge device; When illegal packet capture is detected, the live stream is automatically cut off and an alarm is issued. The live stream fuse mechanism is executed to physically separate the live stream from the conference signaling. Receive meeting record information sent by the edge device, encrypt the meeting record information based on the participant role information, and generate meeting information in the target format; Process the target format meeting information, combine it with timeline-driven speech-to-text conversion and analysis, and generate initial meeting scenario mode information including meeting outline, PPT, and mind map; Based on the priority processing strategy of multimodal information, the target format conference information and the initial conference scene mode information are processed to generate conference information data after priority sorting and strategy processing; Perform semantic analysis, feature fusion, and cross-modal association processing on conference information data to generate target conference scene pattern information, including semantic analysis and feature extraction of conference information data, and semantic parsing of the text content in the conference information data using natural language processing technology to obtain semantic keywords and semantic paragraph information; combine multimodal feature fusion strategies to process semantic keywords, semantic paragraph information, voice features, and video features, perform cross-modal feature association, and generate multimodal fusion feature data; Based on the cross-modal association rules and the logical relationship between conference scenes, the multimodal fusion feature data is structured and integrated to complete the construction of the cross-modal association network; according to the conference scene pattern generation rules and application requirements, the cross-modal association network is verified and abnormal associations are adjusted based on the scene pattern integrity detection mechanism to generate the target conference scene pattern framework; according to the preset scene pattern generation rules, the semantic analysis results, multimodal fusion feature data and cross-modal association network are associated and integrated to generate the target conference scene pattern information, which is returned to the edge device for encrypted storage and backed up locally.

5. The conference live broadcast method according to claim 4, characterized in that: When illegal packet capture is detected, the live stream is automatically disconnected and an alarm is issued. The live stream fuse mechanism is executed to physically separate the live stream from the conference signaling, including: When illegal packet capture is detected, network transmission data is monitored and analyzed in real time to obtain packet capture behavior characteristics and abnormal transmission data; Combine security protection rules with live streaming fuse policies to process packet capture behavior characteristics and abnormal transmission data, conduct security threat assessments, and generate fuse trigger instructions; Based on the physical separation rules of live streaming and conference signaling, the live streaming transmission link and the conference signaling transmission link are isolated to complete the transmission channel separation; Automatically generate alarm information based on live broadcast security requirements and system alarm mechanisms, mark abnormal alarm levels based on the alarm priority detection mechanism, and adjust the sending of alarms; According to the preset circuit breaker execution rules, the illegal packet capture information is associated and encapsulated with the circuit breaker operation instructions to generate an operation data packet with security mark and circuit breaker control, automatically cutting off the live stream and issuing an alarm.

6. The conference live broadcast method according to claim 4, characterized in that: The target format meeting information is processed, combined with timeline-driven speech-to-text conversion and analysis processing to generate initial meeting scenario mode information including meeting outline, PPT, and mind map, including: Combined with timeline-driven speech-to-text conversion and analysis processing, the voice content in the meeting record information is collected and converted in real time to obtain the text information marked on the timeline; Combined with meeting content analysis and keyword extraction strategies, the text information marked on the timeline is processed, semantic segmentation recognition is performed, and preliminary text record data is generated; Based on the timeline nodes, speech paragraph information and meeting outline generation rules, the preliminary text record data is structured and organized to complete the construction of the meeting outline framework; Automatically generate PPT and mind map nodes based on the association of meeting topics and logical relationships of content, and mark abnormal associations and make logical adjustments based on the node association detection mechanism; According to the preset scene mode generation rules, the timeline-driven speech-to-text information is associated and integrated with the meeting outline, PPT, and mind map to generate initial meeting scene mode information including the meeting outline, PPT, and mind map.

7. A conference live broadcast system, characterized in that: include: The edge device receives the HDMI input signal of the video conference computer and the UVC camera input signal, and sends the live broadcast start information to the cloud service device; The cloud service device receives the signal input information and the request to start live broadcast sent by the edge device, generates a time-sensitive QR code with an HMAC signature that automatically expires within a preset time, and sends it to the edge device; when illegal packet capture is detected, the live stream is automatically cut off and an alarm is issued, and the live broadcast fuse mechanism is executed to physically separate the live stream from the conference signaling; the meeting record information sent by the edge device is received, and the meeting record information is encrypted based on the role information of the participants to generate the target format meeting information; the target format meeting information is processed, combined with timeline-driven voice-to-text and analysis processing, to generate the initial meeting scenario mode information including the meeting outline, PPT, and mind map; Based on the priority processing strategy of multimodal information, the target format conference information and the initial conference scene mode information are processed to generate conference information data after priority sorting and strategy processing; Perform semantic analysis, feature fusion, and cross-modal association processing on conference information data to generate target conference scene pattern information, including semantic analysis and feature extraction of conference information data, and semantic parsing of the text content in the conference information data using natural language processing technology to obtain semantic keywords and semantic paragraph information; combine multimodal feature fusion strategies to process semantic keywords, semantic paragraph information, voice features, and video features, perform cross-modal feature association, and generate multimodal fusion feature data; Based on the cross-modal association rules and the logical relationship between conference scenes, the multimodal fusion feature data is structured and integrated to complete the construction of the cross-modal association network. According to the conference scene pattern generation rules and application requirements, the cross-modal association network is verified and abnormal associations are adjusted based on the scene pattern integrity detection mechanism to generate the target conference scene pattern framework. According to the preset scene pattern generation rules, the semantic analysis results, multimodal fusion feature data and cross-modal association network are associated and integrated to generate the target conference scene pattern information, which is returned to the edge device for encrypted storage and backed up locally. The edge device receives a time-sensitive QR code and conference live broadcast information sent by the cloud service device that has an HMAC signature and automatically expires within a preset time; obtains the target user's conference attendance information based on the time-sensitive QR code; processes the target user's conference attendance information to generate conference access rights, target user live broadcast quality information, and a conference record containing the target user's conference speech information; processes the conference live broadcast information based on the conference access rights and target user live broadcast quality information to generate conference live broadcast information; and receives the target conference scene mode information sent by the cloud service device. The target application is confirmed based on the target conference scene mode information, and the target format instruction information is sent to the target application for processing by the target application.

8. A computing device comprising a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein: When the computer program instructions are executed by the processor, the device is triggered to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Intelligent conference multi-mode cooperative control and intelligent processing method

    CN118917819A

  • Method and electronic device for quickly playing video

    US20170171628A1