A resource scheduling method for video conferencing system based on cloud computing

By using cloud computing-based resource scheduling methods in the video conferencing system to record and verify audio and video data, the problem of video conferencing minutes being easily stolen or tampered with is solved, the synchronization and authenticity of meeting minutes are achieved, and all parties participating in the meeting obtain real meeting content.

CN115225849BActive Publication Date: 2025-05-13ANHUI SPIDER INFORMATION & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210839054.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-18
Publication Date
2025-05-13
Estimated Expiration
2042-07-18

AI Technical Summary

Technical Problem

During the video conference, the minutes of the conference are easily illegally stolen or tampered with, resulting in the inability of the parties to the conference to obtain the real content of the conference, posing huge hidden dangers.

Method used

By using cloud computing-based resource scheduling methods in the video conferencing system, the audio and video data of participants from all parties are recorded, and preliminary meeting minutes and identity verification screens are generated locally. Through encrypted transmission and binary overlay processing, the synchronization and authenticity of meeting minutes are ensured.

Benefits of technology

It realizes the timely discovery of meeting minutes being tampered with during the meeting, avoids losses caused by meeting minutes being tampered with, and ensures that all parties participating in the meeting obtain real meeting content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115225849B_ABST
    Figure CN115225849B_ABST
Patent Text Reader

Abstract

The present invention is applicable to the technical field of data transmission, and in particular to a resource scheduling method for a video conference system based on cloud computing, the method comprising: establishing a data connection with a participant's device, and receiving audio and video data uploaded from the participant's device; generating a preliminary meeting minutes, randomly selecting a frame of the picture, and generating a compilation time node; compiling the preliminary meeting minutes into the picture, generating an identity verification picture; receiving the compilation time node, audio and video data, and identity verification picture from an external device, performing picture verification, and determining whether the meeting minutes are synchronized, and if synchronized, storing them. In the present invention, each party obtains audio and video data to generate meeting minutes and transmit them, and by comparing the superimposed images, it is determined whether the meeting minutes obtained by each party are the same, so that during the meeting, it is possible to promptly discover that the meeting minutes have been tampered with, thereby avoiding subsequent losses caused by the tampering of the meeting minutes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data transmission, and in particular relates to a resource scheduling method for a video conferencing system based on cloud computing. Background Art

[0002] Cloud computing is a type of distributed computing, which means breaking down huge data computing programs into countless small programs through the network "cloud", and then processing and analyzing these small programs through a system composed of multiple servers to obtain results and return them to users. In the early days of cloud computing, simply put, it was simple distributed computing, solving task distribution and merging computing results.

[0003] Video conferencing refers to a meeting where people in two or more locations have face-to-face conversations through communication equipment and the Internet. Depending on the number of participating locations, video conferencing can be divided into point-to-point conferencing and multi-point conferencing. Using a video conferencing system, participants can hear the sounds of other venues, see the images, movements and expressions of participants in other venues, and can also send electronic presentation content, making participants feel as if they are actually there.

[0004] In current video conferencing, in order to record the meeting process, the audio or video of the video conference is usually recorded, and the minutes are generated through voice recognition. However, once the minutes of the video conference are illegally stolen or tampered with, the participants will not be able to obtain the real content of the meeting, which will cause huge hidden dangers. Summary of the invention

[0005] The purpose of the embodiment of the present invention is to provide a resource scheduling method for a video conferencing system based on cloud computing, aiming to solve the problem raised in the third part of the background technology.

[0006] The embodiment of the present invention is implemented as follows: a method for scheduling resources of a video conferencing system based on cloud computing, the method comprising:

[0007] Establishing a data connection with a participant's device and receiving audio and video data uploaded from the participant's device, wherein the audio and video data includes timeline data;

[0008] Generate preliminary meeting minutes according to the audio and video data, randomly select a frame from the audio and video data, and generate a compilation time node, where the compilation time node is the time point corresponding to the frame;

[0009] Compile the preliminary meeting minutes into the screen, generate the identity verification screen, and send out the compiled time node, audio and video data, and identity verification screen;

[0010] Receive the compilation time node, audio and video data, and identity verification screen from the external device, perform screen verification, and determine whether the meeting minutes are synchronized. If synchronized, store the preliminary meeting minutes.

[0011] Preferably, the step of generating preliminary meeting minutes according to the audio and video data, randomly selecting a frame from the audio and video data, and generating a compilation time node specifically includes:

[0012] Extract audio and video data to obtain voice data and video data, and identify the speaker based on the video data;

[0013] Determine the speaking device according to the speaker, extract and recognize the corresponding voice data in the speaking device, and generate preliminary meeting minutes;

[0014] A frame is randomly selected from the video data, the corresponding time of the frame is determined, and the compilation time node is obtained.

[0015] Preferably, the steps of compiling the preliminary meeting minutes into the screen, generating an identity verification screen, and sending the compiled time node, audio and video data, and the identity verification screen specifically include:

[0016] Convert the preliminary meeting minutes and the data corresponding to the screen into binary minutes data and binary screen data respectively;

[0017] The binary minutes data and the binary screen data are superimposed to obtain an identity verification screen;

[0018] The compiled time node, audio and video data, and identity verification screen will be sent out.

[0019] Preferably, the step of receiving the compilation time node, audio and video data, and identity verification screen from the external device, performing screen verification, and determining whether the meeting minutes are synchronized specifically includes:

[0020] Receiving a compilation time node, audio and video data, and an identity verification screen from an external device, extracting a frame of screen from the audio and video data from the external device according to the compilation time node, and obtaining an active verification screen;

[0021] Perform speech recognition on the audio and video data from the external device to obtain the minutes of the second meeting;

[0022] The second meeting minutes and the active verification screen are merged in a binary overlay manner to obtain the screen to be verified. The screen to be verified and the identity verification screen are compared to determine whether the meeting minutes are synchronized.

[0023] Preferably, the compilation time node, audio and video data, and identity verification screen are all transmitted in encrypted form.

[0024] Preferably, when the meeting minutes are out of sync, a prompt message is sent to notify all parties to verify the meeting minutes.

[0025] Preferably, at the end of the meeting, final meeting minutes are generated based on the preliminary meeting minutes, and the audio and video data generated during the meeting are uploaded to the cloud for storage.

[0026] Preferably, the audio and video data stored in the cloud are deleted regularly.

[0027] Preferably, the final meeting minutes are sent simultaneously to all participants.

[0028] Another object of an embodiment of the present invention is to provide a video conferencing system resource scheduling system based on cloud computing, the system comprising:

[0029] The device connection module is used to establish a data connection with the device of the participant and receive audio and video data uploaded from the device of the participant, wherein the audio and video data includes timeline data;

[0030] A minutes generation module is used to generate preliminary meeting minutes based on the audio and video data, randomly select a frame from the audio and video data, and generate a compilation time node, where the compilation time node is the time point corresponding to the frame;

[0031] The screen compilation module is used to compile the preliminary meeting minutes into the screen, generate the identity verification screen, and send out the compilation time node, audio and video data, and the identity verification screen;

[0032] The minutes verification module is used to receive the compilation time node, audio and video data, and identity verification images from external devices, perform image verification, and determine whether the meeting minutes are synchronized. If synchronized, the preliminary meeting minutes are stored.

[0033] Preferably, the minutes generation module includes:

[0034] A data extraction unit, used to extract data from the audio and video data to obtain voice data and video data, and identify the speaker based on the video data;

[0035] A speech recognition unit is used to determine the speaking device according to the speaker, extract and recognize the corresponding speech data in the speaking device, and generate preliminary meeting minutes;

[0036] The picture processing unit is used to randomly select a frame of picture from the video data, determine the time corresponding to the picture, and obtain a compilation time node.

[0037] Preferably, the picture compilation module includes:

[0038] A binary conversion unit, used to convert the preliminary meeting minutes and the data corresponding to the screen into binary minutes data and binary screen data respectively;

[0039] A data superposition unit, used for superimposing the binary minutes data and the binary screen data to obtain an identity verification screen;

[0040] The data sending unit is used to send the compilation time node, audio and video data, and identity verification screen.

[0041] Preferably, the minutes verification module includes:

[0042] A data receiving unit, used to receive a compilation time node, audio and video data, and an identity verification screen from an external device, and extract a frame of screen from the audio and video data from the external device according to the compilation time node to obtain an active verification screen;

[0043] A two-speech recognition unit, used to perform speech recognition on the audio and video data from the external device to obtain the second meeting minutes;

[0044] The active verification module is used to merge the minutes of the second meeting with the active verification screen in a binary superposition manner to obtain the screen to be verified, compare the screen to be verified with the identity verification screen, and determine whether the meeting minutes are synchronized.

[0045] A cloud computing-based video conferencing system resource scheduling method provided by an embodiment of the present invention records the audio and video of all participants, thereby locally completing the generation of meeting minutes and images representing the meeting minutes, and then transmitting the audio and video data to all parties. Each party obtains the audio and video data, generates meeting minutes here and transmits them, and determines whether the meeting minutes obtained by all parties are the same by comparing the superimposed images. Therefore, during the meeting, it is possible to promptly discover that the meeting minutes have been tampered with, thereby avoiding subsequent losses caused by the tampering of the meeting minutes. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] Figure 1 A flow chart of a method for scheduling resources of a video conferencing system based on cloud computing provided by an embodiment of the present invention;

[0047] Figure 2 A flowchart of the steps of generating preliminary meeting minutes based on audio and video data, randomly selecting a frame from the audio and video data, and generating a compilation time node provided by an embodiment of the present invention;

[0048] Figure 3 A flowchart of the steps of compiling the preliminary meeting minutes into the screen, generating an identity verification screen, and sending the compiled time node, audio and video data, and the identity verification screen provided by an embodiment of the present invention;

[0049] Figure 4A flowchart of the steps of receiving a compilation time node, audio and video data, and an identity verification screen from an external device, performing screen verification, and determining whether the meeting minutes are synchronized, provided by an embodiment of the present invention;

[0050] Figure 5 An architecture diagram of a resource scheduling system for a video conferencing system based on cloud computing provided by an embodiment of the present invention;

[0051] Figure 6 An architectural diagram of a minutes generation module provided in an embodiment of the present invention;

[0052] Figure 7 An architectural diagram of a picture compilation module provided by an embodiment of the present invention;

[0053] Figure 8 An architectural diagram of a minutes verification module provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0055] It is understood that the terms "first", "second", etc. used in this application may be used herein to describe various elements, but unless otherwise specified, these elements are not limited by these terms. These terms are only used to distinguish a first element from another element. For example, without departing from the scope of this application, a first xx script may be referred to as a second xx script, and similarly, a second xx script may be referred to as a first xx script.

[0056] Video conferencing refers to a meeting in which people in two or more locations have face-to-face conversations through communication equipment and the Internet. Depending on the number of participating locations, video conferencing can be divided into point-to-point conferencing and multi-point conferencing. Using a video conferencing system, participants can hear the sounds of other venues, see the images, movements and expressions of participants in other venues, and can also send electronic presentation content to make participants feel as if they are actually there. In the current video conferencing process, in order to record the meeting process, the audio or video of the video conference is usually recorded, and the minutes of the meeting are generated through voice recognition. However, once the minutes of the video conference are illegally stolen or tampered with, the participants will not be able to obtain the real content of the meeting, which will cause huge hidden dangers.

[0057] In the present invention, by recording the audio and video of the participants from all parties, the meeting minutes and images representing the meeting minutes are generated locally, and then the audio and video data are transmitted to all parties. Each party obtains the audio and video data, generates meeting minutes here and transmits them, and by comparing the superimposed images, it is determined whether the meeting minutes obtained by all parties are the same. Therefore, during the meeting, it is possible to promptly discover that the meeting minutes have been tampered with, thereby avoiding subsequent losses caused by the tampering of the meeting minutes.

[0058] like Figure 1 FIG. 1 is a flowchart of a method for scheduling resources of a video conferencing system based on cloud computing provided by an embodiment of the present invention, the method comprising:

[0059] S100, establishing a data connection with a participant's device and receiving audio and video data uploaded from the participant's device, wherein the audio and video data includes timeline data.

[0060] In this step, a data connection is established with the devices of the participants. The application scenario of the present invention is that multiple parties conduct a video conference, and each party has multiple participants. For example, Company A, Company B and Company C need to conduct a video conference. Company A, Company B and Company C each have multiple employees who need to participate in the conference. Then, when conducting a video conference, Company A, Company B and Company C respectively appoint an employee as a representative. The representatives of the three companies establish a network connection through their respective devices, and other employees of Company A, Company B and Company C are connected to the devices used by the corresponding representatives of their companies. For example, Company A has six employees A1, A2, A3, A4, A5 and A6 participating in the meeting, and A1 is designated as the representative of the company. Then the remaining five employees are connected to the device used by A1 through the device. During the video process, each participant uses the device in his hand to collect audio and video to obtain audio and video data. The device corresponding to the participants of each company transmits the collected audio and video data to the device used by the representative of the company. The audio and video data includes timeline data. The present invention is applied to the devices used by the representatives of each company.

[0061] S200, generating preliminary meeting minutes according to the audio and video data, randomly selecting a frame from the audio and video data, and generating a compilation time node, where the compilation time node is a time point corresponding to the frame.

[0062] In this step, preliminary meeting minutes are generated based on the audio and video data, and the speech information of each participant in the audio and video data is determined through voice recognition. Since the source of the audio and video data is certain, the speech content of each participant can be specifically determined, and the time corresponding to the speech content of each participant can be determined based on the timeline data. In order to ensure that the content of the meeting minutes obtained by all participants is consistent, a frame of the picture is randomly selected from the audio and video data, and the time corresponding to the picture is recorded, that is, the compilation time node is obtained.

[0063] S300, compile the preliminary meeting minutes into the screen, generate an identity verification screen, and send out the compiled time node, audio and video data, and identity verification screen.

[0064] In this step, the preliminary meeting minutes are compiled into the screen. Specifically, the screen data and the content of the preliminary meeting minutes can be unified by base conversion, and then the two data are merged to finally generate an identity verification screen. The compiled time node, audio and video data, and identity verification screen are sent out. Specifically, representatives of each party send the compiled time node, audio and video data, and identity verification screen they generate to other parties; the compiled time node, audio and video data, and identity verification screen are all transmitted in encrypted form.

[0065] S400, receiving the compilation time node, audio and video data and identity verification picture from the external device, performing picture verification, and determining whether the meeting minutes are synchronized. If synchronized, storing the preliminary meeting minutes.

[0066] In this step, the compilation time node, audio and video data, and identity verification screen are received from the external device. The devices of each party compare the identity verification screen to determine whether the meeting minutes obtained by each party are the same. Since the verification is done by checking the pictures, when one party is invaded and tampered, then when the parties compare the identity verification screen, anomalies will occur, and the problem can be discovered in time. Therefore, if they are not synchronized, it means that tampering or anomalies have occurred. If they are synchronized, it means that the content of the meeting minutes currently recorded by all parties is the same, and they can be stored. When the meeting minutes are not synchronized, a prompt message is issued to notify all parties to verify the meeting records. At the end of the meeting, the final meeting minutes are generated based on the preliminary meeting minutes, and the audio and video data generated during the meeting are uploaded to the cloud for storage. The audio and video data stored in the cloud are deleted regularly. The final meeting minutes are sent synchronously to all participants.

[0067] like Figure 2 As shown, as a preferred embodiment of the present invention, the steps of generating preliminary meeting minutes according to audio and video data, randomly selecting a frame from the audio and video data, and generating a compilation time node specifically include:

[0068] S201, extracting audio and video data to obtain voice data and video data, and identifying a speaker based on the video data.

[0069] In this step, data extraction is performed on the audio and video data, the audio tracks in the audio and video data are separated, and the picture data is separately stripped out to obtain voice data and video data. Through portrait recognition, it is determined whether the persons in each video data are speaking. For example, if two participants attend the meeting at the same time, due to the close distance between them, the two devices may record the sound of one of them speaking at the same time, which will cause inaccuracies when generating the meeting minutes. In order to solve this problem, the video data of each device can be used to determine whether the user corresponding to the device is speaking, thereby determining the source of the voice corresponding to the video segment to determine the speaker. Of course, the device can also analyze the video when recording audio to determine whether the user has mouth movements. If there are no mouth movements, no recording will be performed, and recording will only start when there are mouth movements.

[0070] S202, determining a speaking device according to the speaker, extracting and identifying corresponding voice data in the speaking device, and generating preliminary meeting minutes.

[0071] In this step, the speaker and the speaking device are determined according to the speaker, and speech recognition is completed through the speech recognition engine. According to the speaking order, preliminary conversational meeting minutes are generated to facilitate subsequent verification.

[0072] S203, randomly selecting a frame from the video data, determining the time corresponding to the frame, and obtaining a compilation time node.

[0073] In this step, a frame is randomly selected from the video data. Specifically, a random function can be used to generate a time, and then the frame corresponding to the time is determined based on the time. Therefore, when the frame is obtained, the compilation time node can be obtained.

[0074] like Figure 3 As shown, as a preferred embodiment of the present invention, the steps of compiling the preliminary meeting minutes into the screen, generating an identity verification screen, and sending the compiled time node, audio and video data, and identity verification screen specifically include:

[0075] S301, converting the preliminary meeting minutes and the data corresponding to the screen into binary minutes data and binary screen data respectively.

[0076] In this step, base conversion is performed, and the preliminary meeting minutes are represented in binary. The data corresponding to the same screen is also represented in binary. At this time, the expression method between the two is the same, which is convenient for processing.

[0077] S302, superimposing the binary minutes data and the binary screen data to obtain an identity verification screen.

[0078] In this step, the binary minutes data and the binary screen data are superimposed. The binary minutes data must contain less data content than the binary screen data. Therefore, when superimposing, they are superimposed as binary numbers. For example, if the binary minutes data is 1100 and the binary screen data is 100101010110, the binary screen data is divided into multiple binary character strings with the same length as the binary minutes data, and then addition is performed, such as 1001+1100 to obtain 10101. Characters exceeding its length are discarded to obtain 0101. Finally, 011000100010 is obtained according to the calculation, and it is converted into the identity verification screen.

[0079] S303, the compilation time node, audio and video data and identity verification screen are sent out.

[0080] In this step, the compilation time node, audio and video data, and identity verification screen are sent out, that is, the representative device that sends the compilation time node, audio and video data, and identity verification screen is generated and sent to the representative devices of other parties for verification by other devices.

[0081] like Figure 4 As shown, as a preferred embodiment of the present invention, the steps of receiving the compilation time node, audio and video data, and identity verification screen from the external device, performing screen verification, and determining whether the meeting minutes are synchronized specifically include:

[0082] S401, receiving a compilation time node, audio and video data, and an identity verification picture from an external device, extracting a frame of picture from the audio and video data from the external device according to the compilation time node, and obtaining an active verification picture.

[0083] In this step, the compilation time node, audio and video data, and identity verification screen are received from the external device, and the screen is extracted according to the compilation time node to obtain a frame of screen. Since the audio and video data is sent by its generator, if there is no tampering, the extracted screen should be the same as the screen corresponding to the corresponding moment of the sender.

[0084] S402, performing voice recognition on the audio and video data from the external device to obtain the minutes of the second meeting.

[0085] In this step, voice recognition is performed on the audio and video data from the external device, and the same voice recognition engine is used for recognition to avoid content deviation caused by different voice recognition engines, and the minutes of the second meeting are obtained.

[0086] S403, the second meeting minutes and the active verification screen are merged in a binary superposition manner to obtain a to-be-verified screen, and the to-be-verified screen is compared with the identity verification screen to determine whether the meeting minutes are synchronized.

[0087] In this step, the same overlay method is used to convert it into binary, and then overlay it. Then, through the conversion, the picture to be verified is obtained. Then, through pixel comparison, the picture to be verified and the identity verification picture are compared. If the two are the same, it means that the contents of the meeting minutes of all parties are consistent.

[0088] like Figure 5 As shown, a video conferencing system resource scheduling system based on cloud computing is provided in an embodiment of the present invention, and the system includes:

[0089] The device connection module 100 is used to establish a data connection with the device of the participant and receive audio and video data uploaded from the device of the participant, wherein the audio and video data includes timeline data.

[0090] In this system, the device connection module 100 establishes a data connection with the device of the participant. The application scenario of the present invention is that multiple parties conduct a video conference, and each party has multiple participants. For example, Company A, Company B and Company C need to conduct a video conference. Company A, Company B and Company C each have multiple employees who need to participate in the conference. Then, when conducting a video conference, Company A, Company B and Company C respectively appoint an employee as a representative. The representatives of the three companies establish a network connection through their respective devices, and other employees of Company A, Company B and Company C are connected to the devices used by the corresponding representatives of their respective companies. For example, Company A has six employees A1, A2, A3, A4, A5 and A6 participating in the conference, and A1 is designated as the representative of the company. Then the remaining five employees are connected to the device used by A1 through the device. During the video process, each participant uses the device in his hand to collect audio and video to obtain audio and video data. The device corresponding to the participant of each company transmits the collected audio and video data to the device used by the representative of the company. The audio and video data includes timeline data. The present invention is applied to the devices used by the representatives of each company.

[0091] The minutes generation module 200 is used to generate preliminary meeting minutes based on the audio and video data, randomly select a frame from the audio and video data, and generate a compilation time node, where the compilation time node is the time point corresponding to the frame.

[0092] In this system, the minutes generation module 200 generates preliminary meeting minutes based on the audio and video data, and determines the speech information of each participant in the audio and video data through voice recognition. Since the source of the audio and video data is certain, the speech content of each participant can be specifically determined, and the time corresponding to the speech content of each participant can be determined based on the timeline data. In order to ensure that the content of the meeting minutes obtained by all participants is consistent, a frame of the picture is randomly selected from the audio and video data, and the time corresponding to the picture is recorded, that is, the compilation time node is obtained.

[0093] The screen compilation module 300 is used to compile the preliminary meeting minutes into the screen, generate an identity verification screen, and send out the compiled time node, audio and video data, and the identity verification screen.

[0094] In this system, the screen compilation module 300 compiles the preliminary meeting minutes into the screen. Specifically, the screen data and the content of the preliminary meeting minutes can be unified by base conversion, and then the data of the two can be merged to finally generate an identity verification screen, and the compiled time node, audio and video data, and identity verification screen are sent out. Specifically, representatives of each party send the compiled time node, audio and video data, and identity verification screen they generate to other parties; the compiled time node, audio and video data, and identity verification screen are all transmitted in encrypted form.

[0095] The minutes verification module 400 is used to receive the compilation time node, audio and video data and identity verification images from the external device, perform image verification, and determine whether the meeting minutes are synchronized. If synchronized, the preliminary meeting minutes are stored.

[0096] In this system, the minutes verification module 400 receives the compilation time node, audio and video data, and identity verification screen from the external device. The devices of each party compare the identity verification screen to determine whether the meeting minutes obtained by each party are the same. Since the verification is done by pictures, when one party is invaded and tampered with, then when the parties compare the identity verification screens, anomalies will occur, and problems can be discovered in time. Therefore, if they are not synchronized, it means that tampering or anomalies have occurred. If they are synchronized, it means that the content of the meeting minutes currently recorded by each party is the same, and it can be stored.

[0097] like Figure 6 As shown, as a preferred embodiment of the present invention, the minutes generation module 200 includes:

[0098] The data extraction unit 201 is used to extract data from the audio and video data to obtain voice data and video data, and identify the speaker according to the video data.

[0099] In this module, the data extraction unit 201 extracts data from the audio and video data, separates the audio track in the audio and video data, and separately strips out the picture data to obtain voice data and video data. By means of portrait recognition, it is determined whether the persons in each video data are speaking. For example, if two participants attend the meeting at the same time, due to the close distance between them, the two devices may simultaneously record the sound of one of them speaking, which will cause inaccuracies when generating the meeting minutes. In order to solve this problem, it is possible to use the video data of each device to determine whether the user corresponding to the device is speaking, thereby determining the source of the voice corresponding to the video segment to determine the speaker. Of course, the device can also analyze the audio based on the video when recording, and determine whether the user has mouth movements. If there are no mouth movements, no recording will be performed, and recording will only start when there are mouth movements.

[0100] The speech recognition unit 202 is used to determine the speaking device according to the speaker, extract and recognize the corresponding speech data in the speaking device, and generate preliminary meeting minutes.

[0101] In this module, the speech recognition unit 202 determines the speaker and the speaking device according to the speaker, completes speech recognition through the speech recognition engine, and generates a conversational preliminary meeting minutes according to the speaking order for subsequent verification.

[0102] The picture processing unit 203 is used to randomly select a frame of picture from the video data, determine the time corresponding to the picture, and obtain a compilation time node.

[0103] In this module, the picture processing unit 203 randomly selects a frame from the video data. Specifically, a random function can be used to generate a time, and then the picture corresponding to the time is determined according to the time. Therefore, when the picture is obtained, the compilation time node can be obtained.

[0104] like Figure 7 As shown, as a preferred embodiment of the present invention, the picture compilation module 300 includes:

[0105] The binary conversion unit 301 is used to convert the preliminary meeting minutes and the data corresponding to the screen into binary minutes data and binary screen data respectively.

[0106] In this module, the base conversion unit 301 performs base conversion and represents the preliminary meeting minutes in binary. The data corresponding to the same screen is also represented in binary. At this time, the expression methods of the two are the same, which is convenient for processing.

[0107] The data superposition unit 302 is used to superimpose the binary minutes data and the binary screen data to obtain an identity verification screen.

[0108] In this module, the data superposition unit 302 superimposes the binary minutes data and the binary screen data. For the binary minutes data, the data content it contains is necessarily less than the binary screen data. Therefore, when superimposing, it is superimposed as a binary number. For example, if the binary minutes data is 1100 and the binary screen data is 100101010110, the binary screen data is divided into multiple binary character strings with the same length as the binary minutes data, and then addition is performed, such as 1001+1100 to obtain 10101. The characters exceeding its length are discarded to obtain 0101. Finally, 011000100010 is obtained according to the calculation, and it is converted into the identity verification screen.

[0109] The data sending unit 303 is used to send the compilation time node, audio and video data, and identity verification screen.

[0110] In this module, the data sending unit 303 sends the compilation time node, audio and video data, and identity verification screen, that is, generates the above-mentioned representative device that sends the compilation time node, audio and video data, and identity verification screen, and sends it to the representative devices of other parties for verification by other devices.

[0111] like Figure 8 As shown, as a preferred embodiment of the present invention, the minutes verification module 400 includes:

[0112] The data receiving unit 401 is used to receive the compilation time node, audio and video data and identity verification picture from the external device, extract a frame of picture from the audio and video data from the external device according to the compilation time node, and obtain the active verification picture.

[0113] In this module, the data receiving unit 401 receives the compilation time node, audio and video data, and identity verification screen from an external device, extracts the screen according to the compilation time node, and obtains a frame of screen. Since the audio and video data is sent by its generator, if there is no tampering, the extracted screen should be the same as the screen corresponding to the corresponding moment of the sender.

[0114] The second audio recognition unit 402 is used to perform speech recognition on the audio and video data from the external device to obtain the second meeting minutes.

[0115] In this module, the two-speech recognition unit 402 performs speech recognition on the audio and video data from the external device, and uses the same speech recognition engine to avoid content deviation caused by different speech recognition engines, and obtains the second meeting minutes.

[0116] The active verification module 403 is used to merge the two meeting minutes and the active verification screen in a binary superposition manner to obtain the screen to be verified, compare the screen to be verified with the identity verification screen, and determine whether the meeting minutes are synchronized.

[0117] In this module, the active verification module 403 adopts the same superposition method to convert it into binary, and then superimposes it. Then, through the conversion, the picture to be verified is obtained, and then the picture to be verified and the identity verification picture are compared through pixel comparison. If the two are the same, it means that the contents of the meeting minutes of all parties are consistent.

[0118] It should be understood that, although each step in the flow chart of each embodiment of the present invention is shown in sequence according to the indication of the arrow, these steps are not necessarily performed in sequence according to the order indicated by the arrow. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0119] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0120] The technical features of the above-described embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above-described embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0121] The above-mentioned embodiments only express several implementation methods of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the patent of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

[0122] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for scheduling resources of a video conferencing system based on cloud computing, characterized in that: The method comprises: Establishing a data connection with a participant's device and receiving audio and video data uploaded from the participant's device, wherein the audio and video data includes timeline data; Generate preliminary meeting minutes according to the audio and video data, randomly select a frame from the audio and video data, and generate a compilation time node, where the compilation time node is the time point corresponding to the frame; Compile the preliminary meeting minutes into the screen, generate the identity verification screen, and send out the compiled time node, audio and video data, and identity verification screen; Receive the compilation time node, audio and video data, and identity verification screen from the external device, perform screen verification, and determine whether the meeting minutes are synchronized. If synchronized, store the preliminary meeting minutes; The steps of compiling the preliminary meeting minutes into the screen, generating an identity verification screen, and sending the compiled time node, audio and video data, and the identity verification screen specifically include: Convert the preliminary meeting minutes and the data corresponding to the screen into binary minutes data and binary screen data respectively; The binary minutes data and the binary screen data are superimposed to obtain an identity verification screen; Send out the compilation time node, audio and video data, and identity verification screen; The step of receiving the compilation time node, audio and video data, and identity verification screen from the external device, performing screen verification, and determining whether the meeting minutes are synchronized specifically includes: Receiving a compilation time node, audio and video data, and an identity verification screen from an external device, extracting a frame of screen from the audio and video data from the external device according to the compilation time node, and obtaining an active verification screen; Perform speech recognition on the audio and video data from the external device to obtain the minutes of the second meeting; The second meeting minutes and the active verification screen are merged in a binary overlay manner to obtain the screen to be verified. The screen to be verified and the identity verification screen are compared to determine whether the meeting minutes are synchronized.

2. The method for scheduling resources of a video conference system based on cloud computing according to claim 1, characterized in that: The step of generating preliminary meeting minutes according to the audio and video data, randomly selecting a frame from the audio and video data, and generating a compilation time node specifically includes: Extract the audio and video data to obtain voice data and video data, and identify the speaker based on the video data; Determine the speaking device according to the speaker, extract and recognize the corresponding voice data in the speaking device, and generate preliminary meeting minutes; A frame is randomly selected from the video data, the corresponding time of the frame is determined, and the compilation time node is obtained.

3. The method for scheduling resources of a video conferencing system based on cloud computing according to claim 1, characterized in that: The compilation time node, audio and video data, and identity verification screen are all transmitted in encrypted form.

4. The method for scheduling resources of a video conference system based on cloud computing according to claim 1, characterized in that: When the meeting minutes are out of sync, a reminder message is sent to notify all parties to verify the meeting minutes.

5. The method for scheduling resources of a video conference system based on cloud computing according to claim 1, characterized in that: At the end of the meeting, final meeting minutes are generated based on the preliminary meeting minutes, and the audio and video data generated during the meeting are uploaded to the cloud for storage.

6. The method for scheduling resources of a video conference system based on cloud computing according to claim 5, characterized in that: Audio and video data stored in the cloud are deleted regularly.

7. The method for scheduling resources of a video conference system based on cloud computing according to claim 5, characterized in that: The final meeting minutes are sent simultaneously to all participants.

Citation Information

Patent Citations

  • Conference sharing method and device and conference record generating method and device

    CN107911646A

  • Data processing method and device, equipment and storage medium

    CN113014540A