A method, device, system and readable storage medium for sending virtual business cards

By receiving audio and video streaming data on the server side of remote video conferencing, determining spokespersons and generating virtual business cards, the problem of low business card sharing efficiency in remote video conferencing is solved, and communication efficiency and user experience are improved.

CN114868380BActive Publication Date: 2025-06-10BOE TECHNOLOGY GROUP CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202080002943.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-23
Publication Date
2025-06-10
Estimated Expiration
2040-11-23

AI Technical Summary

Technical Problem

In remote video conferencing, more participants lead to low efficiency in business card sharing, which in turn affects communication efficiency and user experience.

Method used

By receiving audio and video stream data on the server, determining the spokesperson and obtaining their biometric information, generating a virtual business card, superimposing it into the audio and video stream data, synthesizing the audio and video stream data to be sent, and sending it to the participating terminal.

Benefits of technology

Improve the efficiency of business card sharing in remote video conferencing, enhance communication efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114868380B_ABST
    Figure CN114868380B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method, apparatus, system and readable storage medium for sending virtual business cards, receiving at least one audio-video stream data of a plurality of participating terminals communicatively connected to a server (S101); determining, from the at least one audio-video stream data, target audio-video stream data corresponding to at least one speaker (S102); obtaining, from the target audio-video stream data, biometric information for identifying the at least one speaker, where the biometric information includes at least one of face feature information and voiceprint feature information (S103); generating virtual business cards for the at least one speaker respectively according to the biometric information of the at least one speaker respectively (S104); superimposing the virtual business cards of the at least one speaker respectively onto the target audio-video stream data, and synthesizing the target audio-video stream data with other audio-video stream data except the target audio-video stream data in the plurality of audio-video stream data into a to-be-sent audio-video stream data (S105); and sending the to-be-sent audio-video stream data to the plurality of participating terminals (S106).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of information technology, and particularly to a method, apparatus, system, and readable storage medium for sending virtual business cards. Background Art

[0002] With the development of information technology, due to the high efficiency of meeting communication, remote video conferencing has gradually replaced traditional round-table meetings.

[0003] However, in remote video conferencing, a large number of participants will inevitably lead to low efficiency of business card sharing, which in turn leads to low communication efficiency and poor user experience. Summary of the Invention

[0004] The present disclosure provides a method, apparatus, system, and readable storage medium for sending virtual business cards. The specific solutions are as follows:

[0005] An embodiment of the present disclosure provides a method for sending virtual business cards, which is applied to a server. The method includes:

[0006] Receiving at least one audio-video stream data of a plurality of participating terminals communicatively connected to the server;

[0007] Determining, from the at least one audio-video stream data, target audio-video stream data corresponding to at least one speaker;

[0008] Obtaining, from the target audio-video stream data, biometric information for identifying each of the at least one speaker, where the biometric information includes at least one of face feature information and voiceprint feature information;

[0009] Generating, according to the biometric information of each of the at least one speaker, a virtual business card for each of the at least one speaker;

[0010] Overlaying the virtual business card of each of the at least one speaker on the target audio-video stream data, and synthesizing the target audio-video stream data with other audio-video stream data in the plurality of audio-video stream data into a to-be-sent audio-video stream data;

[0011] Sending the to-be-sent audio-video stream data to the plurality of participating terminals, so that the plurality of participating terminals display the virtual business card of each of the at least one speaker.

[0012] Optionally, in an embodiment of the present disclosure, if a plurality of different voiceprint feature information are simultaneously obtained from the target audio-video stream data, the method further includes:

[0013] Determining that the at least one speaker is a plurality of speakers corresponding to the plurality of different voiceprint feature information.

[0014] Optionally, in the embodiments of the present disclosure, if multiple different voiceprint feature information are successively obtained from the target audio-video stream data within a preset duration, the method further includes:

[0015] Determine that the at least one speaker is the multiple speakers corresponding to the multiple different feature information.

[0016] Optionally, in the embodiments of the present disclosure, the overlaying the respective virtual business cards of the at least one speaker onto the target audio-video stream data includes:

[0017] For each of the at least one speaker, detect the coordinate position of the face region of the corresponding speaker from the target audio-video stream data;

[0018] Determine the target position and size of the virtual business card of the speaker according to the coordinate position of the face region of the speaker;

[0019] Overlay the virtual business card onto the target audio-video stream data according to the target position and size of the virtual business card.

[0020] Optionally, in the embodiments of the present disclosure, if the face region of the at least one speaker is not detected in the target audio-video stream data, the method further includes:

[0021] Overlay the respective virtual business cards of the at least one speaker onto the target audio-video stream data according to the preset coordinate positions.

[0022] Optionally, in the embodiments of the present disclosure, the overlaying the respective virtual business cards of the at least one speaker onto the target audio-video stream data includes:

[0023] Calculate the average gray scale value corresponding to each color channel in a preset color channel of at least one image in the target audio-video stream data, and adjust the chromaticity of the virtual business card according to the ratio of the average gray scale values corresponding to the respective color channels, so as to obtain the adjusted virtual business cards of the at least one speaker respectively, so that the contrast between the chromaticity of the adjusted virtual business cards of the at least one speaker and the chromaticity of the at least one image is greater than a preset value;

[0024] Overlay the adjusted virtual business cards of the at least one speaker onto the target audio-video stream data.

[0025] Optionally, in the embodiments of the present disclosure, the determining the target position and size of the virtual business card of the speaker according to the coordinate position of the face region of the speaker includes:

[0026] Determine the coordinate positions of other regions in the target audio-video stream data except for the face region of the speaker according to the coordinate positions of the face region of the speaker;

[0027] Determine at least one connected region from the other regions according to the coordinate positions of the other regions;

[0028] Determine a target connected region with an area larger than a preset area from the at least one connected region, and the coordinate positions of the target connected region;

[0029] Determine the target position and size of the virtual business card of the speaker according to the target connected region and the coordinate positions of the target connected region.

[0030] Optionally, in the embodiments of the present disclosure, the determining the target position and size of the virtual business card of the speaker according to the target connected region and the coordinate positions of the target connected region includes:

[0031] Determine the region with the largest area in the target connected region that is the same as the preset business card shape according to the preset business card shape, and adjust the target position and size of the virtual business card of the speaker according to the coordinate positions where the region with the largest area is located.

[0032] Optionally, in the embodiments of the present disclosure, the determining the target position and size of the virtual business card of the speaker according to the target connected region and the coordinate positions of the target connected region includes:

[0033] Determine the inscribed figure with the largest area within the target connected region;

[0034] Take the inscribed figure with the largest area as the shape of the virtual business card of the speaker;

[0035] Adjust the target position and size of the virtual business card of the speaker according to the coordinate positions where the inscribed figure with the largest area is located.

[0036] Optionally, in the embodiments of the present disclosure, the overlaying the virtual business cards of each of the at least one speaker onto the target audio-video stream data and synthesizing the virtual business cards with other audio-video stream data except the target audio-video stream data in the plurality of audio-video stream data into a to-be-transmitted audio-video stream data includes:

[0037] According to a preset layout, the target audio-video stream data with the virtual business cards of at least one speaker superimposed thereon is synthesized with the other audio-video stream data in the multiple audio-video stream data except the target audio-video stream data to form a to-be-sent audio-video stream data, so that the video picture corresponding to the target audio-video stream data with the virtual business cards of at least one speaker superimposed thereon among the multiple participating terminals is larger than the video pictures corresponding to the other audio-video stream data.

[0038] Optionally, in an embodiment of the present disclosure, before receiving the multiple audio-video stream data of multiple participating terminals in the received communication connection, the method further includes:

[0039] Pre-storing the corresponding relationship between the biometric information of known participants and the virtual business cards.

[0040] Optionally, in an embodiment of the present disclosure, if the biometric information of at least one speaker is not obtained in the target audio-video stream data, the method further includes:

[0041] Determining the video pictures of at least one speaker respectively;

[0042] Receiving a screenshot operation of a person with input permission for the video pictures of at least one speaker respectively, and in response to the screenshot operation, determining the biometric information of at least one speaker respectively;

[0043] Receiving a text input operation for the video pictures of at least one speaker respectively, and in response to the text input operation, determining the virtual business cards of at least one speaker respectively;

[0044] Associating the biometric information of at least one speaker respectively and the virtual business cards of at least one speaker respectively.

[0045] Optionally, in an embodiment of the present disclosure, if no speaker is detected in the multiple audio-video stream data, the method further includes:

[0046] Determining the conference terminal corresponding to the conference host from the multiple participating terminals, and using the audio-video stream data corresponding to the conference terminal as the target audio-video stream data.

[0047] Correspondingly, an embodiment of the present disclosure provides a virtual business card sending device, which is applied to a server side and includes:

[0048] A receiving unit, configured to receive at least one audio-video stream data of multiple participating terminals communicatively connected to the server side;

[0049] A determination unit, configured to determine, from the at least one audio-video stream data, target audio-video stream data corresponding to at least one speaker;

[0050] An acquisition unit, configured to acquire, from the target audio-video stream data, biometric information for identifying each of the at least one speaker, where the biometric information includes at least one of face feature information and voiceprint feature information;

[0051] A generation unit, configured to generate, according to the biometric information of each of the at least one speaker, virtual business cards for each of the at least one speaker;

[0052] A synthesis unit, configured to superimpose the virtual business cards of each of the at least one speaker onto the target audio-video stream data, and synthesize the virtual business cards with other audio-video stream data except the target audio-video stream data in the multiple audio-video stream data into a to-be-transmitted audio-video stream data;

[0053] A transmission unit, configured to transmit the to-be-transmitted audio-video stream data to the multiple participating terminals, so that the multiple participating terminals display the virtual business cards of each of the at least one speaker.

[0054] Optionally, in an embodiment of the present disclosure, if the acquisition unit simultaneously acquires multiple different voiceprint feature information from the target audio-video stream data, the determination unit is further configured to:

[0055] Determine that the at least one speaker is multiple speakers corresponding to the multiple different voiceprint feature information.

[0056] Optionally, in an embodiment of the present disclosure, if the acquisition unit sequentially acquires multiple different voiceprint feature information from the target audio-video stream data, the determination unit is further configured to:

[0057] Determine that the at least one speaker is multiple speakers corresponding to the multiple different feature information.

[0058] Optionally, in an embodiment of the present disclosure, the synthesis unit is configured to:

[0059] For each of the at least one speaker, detect the coordinate position of the face area of the corresponding speaker from the target audio-video stream data;

[0060] According to the coordinate position of the face area of the speaker, determine the target position and size of the virtual business card of the speaker;

[0061] Overlay the virtual business card of the speaker onto the target audio-video stream data according to the target position and size of the virtual business card. Optionally, in the embodiments of the present disclosure, if no face region of the speaker is detected in the target audio-video stream data, the synthesizing unit is further configured to:

[0062] Overlay the virtual business cards of the at least one speaker onto the target audio-video stream data according to preset coordinate positions.

[0063] Optionally, in the embodiments of the present disclosure, the synthesizing unit is configured to:

[0064] Calculate the average gray-scale value corresponding to each color channel in a preset color channel of at least one image in the target audio-video stream data, and adjust the chroma of the virtual business cards of the at least one speaker according to the ratio of the average gray-scale values corresponding to the respective color channels, so as to obtain the adjusted virtual business cards of the at least one speaker, such that the contrast between the chroma of the adjusted virtual business cards of the at least one speaker and the chroma of the at least one image is greater than a preset value;

[0065] Overlay the adjusted virtual business cards of the at least one speaker onto the target audio-video stream data.

[0066] Optionally, in the embodiments of the present disclosure, the synthesizing unit is configured to:

[0067] Determine the coordinate positions of other regions in the target audio-video stream data except for the face region of the speaker according to the coordinate positions of the face region of the speaker;

[0068] Determine at least one connected region from the other regions according to the coordinate positions of the other regions;

[0069] Determine a target connected region with an area greater than a preset area and the coordinate positions of the target connected region from the at least one connected region;

[0070] Determine the target position and size of the virtual business card of the speaker according to the target connected region and the position of the target connected region.

[0071] Optionally, in the embodiments of the present disclosure, the synthesizing unit is configured to:

[0072] Determine the region with the largest area having the same shape as the preset business card shape in the target connected region, and adjust the target position and size of the virtual business card of the speaker according to the coordinate position where the region with the largest area is located.

[0073] Optionally, in the embodiments of the present disclosure, the synthesizing unit is configured to:

[0074] Determine the inscribed figure with the largest area within the target connected region;

[0075] Use the inscribed figure with the largest area as the shape of the virtual business card of the speaker;

[0076] Adjust the target position and size of the virtual business card of the speaker according to the coordinate position where the inscribed figure with the largest area is located.

[0077] Optionally, in the embodiments of the present disclosure, the synthesis unit is configured to:

[0078] According to a preset layout, synthesize the target audio-video stream data with the virtual business cards of the at least one speaker superimposed thereon and the other audio-video stream data in the multiple audio-video stream data except the target audio-video stream data into a to-be-sent audio-video stream data, so that the video picture corresponding to the target audio-video stream data with the virtual business cards of the at least one speaker superimposed thereon in the multiple participating terminals is larger than the video picture corresponding to the other audio-video stream data.

[0079] Optionally, in the embodiments of the present disclosure, the device further includes a storage unit, and the storage unit is configured to:

[0080] Pre-store the correspondence between the biometric information of known participants and virtual business cards.

[0081] Optionally, in the embodiments of the present disclosure, if the biometric information of the at least one speaker is not obtained in the target audio-video stream data, the device further includes an input unit, and the input unit is configured to:

[0082] Determine the video pictures of the at least one speaker respectively;

[0083] Receive the screenshot operation of the person with input permission for the video pictures of the at least one speaker respectively, and in response to the screenshot operation, determine the biometric information of the at least one speaker respectively;

[0084] Receive the text input operation for the video pictures of the at least one speaker respectively, and in response to the text input operation, determine the virtual business cards of the at least one speaker respectively;

[0085] Associate the biometric information of the at least one speaker respectively and the virtual business cards of the at least one speaker respectively. Correspondingly, the embodiments of the present disclosure provide a virtual business card sending system, wherein the sending system includes a server and multiple participating terminals communicatively connected to the server;

[0086] The server is configured to receive at least one audio - video stream data of a plurality of participating terminals communicatively connected to the server;

[0087] The server is further configured to determine, from the at least one audio - video stream data, target audio - video stream data corresponding to the at least one speaker;

[0088] The server is further configured to obtain, from the target audio - video stream data, biometric information for identifying each of the at least one speaker, where the biometric information includes at least one of face feature information and voiceprint feature information;

[0089] The server is further configured to generate virtual business cards for each of the at least one speaker according to the biometric information of each of the at least one speaker; superimpose the virtual business cards of each of the at least one speaker on the target audio - video stream data, and synthesize the target audio - video stream data with other audio - video stream data in the plurality of audio - video stream data into a to - be - sent audio - video stream data; and send the to - be - sent audio - video stream data to the plurality of participating terminals;

[0090] The plurality of participating terminals are configured to display the virtual business cards of each of the at least one speaker.

[0091] Optionally, in an embodiment of the present disclosure, the server includes a conference server and a feature recognition server communicatively connected to the conference server, where the conference server is configured to receive the plurality of audio - video stream data, determine the target audio - video stream data from the plurality of audio - video stream data, and determine virtual business cards corresponding to the biometric information of each of the at least one speaker by invoking a feature detection and recognition interface of the feature recognition server, and send the to - be - sent audio - video stream data with the virtual business cards superimposed thereon to the plurality of participating terminals;

[0092] The feature recognition server is configured to identify, from the target audio - video stream data, biometric information for identifying each of the at least one speaker, and send the biometric information of each of the at least one speaker to the conference server.

[0093] Optionally, in an embodiment of the present disclosure, the plurality of participating terminals are further configured to:

[0094] Display the video picture corresponding to the target audio - video stream data with the virtual business card superimposed thereon in a size larger than the video pictures corresponding to other audio - video stream data in the plurality of audio - video stream data except the target audio stream data.

[0095] Correspondingly, an embodiment of the present disclosure provides a virtual business card sending device, which includes:

[0096] A memory and a processor;

[0097] Wherein, the memory is used to store a computer program;

[0098] The processor is used to execute the computer program in the memory to implement the following steps:

[0099] Receiving at least one audio - video stream data of multiple participating terminals communicatively connected to a server;

[0100] Determining, from the at least one audio - video stream data, target audio - video stream data corresponding to at least one speaker;

[0101] Obtaining, from the target audio - video stream data, biometric information for identifying each of the at least one speaker, where the biometric information includes at least one of face feature information and voiceprint feature information;

[0102] Generating respective virtual business cards for each of the at least one speaker according to the biometric information of each of the at least one speaker;

[0103] Overlaying the respective virtual business cards of the at least one speaker onto the target audio - video stream data, and synthesizing the target audio - video stream data with other audio - video stream data in the multiple audio - video stream data except the target audio - video stream data into a to - be - sent audio - video stream data;

[0104] Sending the to - be - sent audio - video stream data to the multiple participating terminals, so that the multiple participating terminals display the respective virtual business cards of the at least one speaker. Correspondingly, an embodiment of the present disclosure provides a computer non - transient readable storage medium, wherein:

[0105] The storage medium stores computer instructions, and when the computer instructions run on a computer, the computer is made to execute the method for sending a virtual business card as described above. Description of the Drawings

[0106] Figure 1 It is a flowchart of a method for sending a virtual business card provided by an embodiment of the present disclosure;

[0107] Figure 2 It is one of the flowcharts of step S105 in the method for sending a virtual business card provided by an embodiment of the present disclosure;

[0108] Figure 3 It is one of the flowcharts of step S105 in the method for sending a virtual business card provided by an embodiment of the present disclosure;

[0109] Figure 4The flowchart of step S202 in a method for sending a virtual business card provided by an embodiment of the present disclosure;

[0110] Figure 5 The flowchart of the second implementation manner of step S404 in a method for sending a virtual business card provided by an embodiment of the present disclosure;

[0111] Figure 6 One of the schematic diagrams of a preset layout in a method for sending a virtual business card provided by an embodiment of the present disclosure;

[0112] Figure 7 One of the schematic diagrams of a preset layout in a method for sending a virtual business card provided by an embodiment of the present disclosure;

[0113] Figure 8 The flowchart of a method for a virtual business card sending method provided by an embodiment of the present disclosure when biometric information of at least one speaker is not obtained in target audio - video stream data;

[0114] Figure 9 One of the structural block diagrams of a device for sending a virtual business card provided by an embodiment of the present disclosure;

[0115] Figure 10 One of the structural block diagrams of a system for sending a virtual business card provided by an embodiment of the present disclosure;

[0116] Figure 11 One of the structural block diagrams of a device for sending a virtual business card provided by an embodiment of the present disclosure. Detailed implementation manners

[0117] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. And without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts fall within the scope of protection of the present disclosure.

[0118] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The terms such as "including" or "comprising" used in the present disclosure mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects.

[0119] In the prior art, there is a technical problem of low efficiency in business card sharing in remote video conferencing.

[0120] In view of this, embodiments of the present disclosure provide a method, apparatus, system, and readable storage medium for sending virtual business cards, which are used to improve the efficiency of business card sharing in remote video conferencing.

[0121] As Figure 1 shown, embodiments of the present disclosure provide a method for sending virtual business cards, which is applied to a server, and includes:

[0122] S101: Receive at least one audio-video stream data of a plurality of participating terminals communicatively connected to the server;

[0123] In a specific implementation process, the server includes a conference server and a feature recognition server communicatively connected to the conference server. The conference server includes a streaming media service module and a conference management service module. The streaming media service module is used to process audio-video stream data, and can be used for audio-video encoding and decoding, invoking a face recognition interface for face recognition, video image overlay processing, audio-video real-time communication, etc. The conference management service module is used to process video conference services, such as participating personnel management, appointment of meetings, joining meetings, meeting notifications, meeting control, sharing and collaboration, background management, etc. The feature recognition server can deploy face detection and face recognition algorithms to perform face detection on the received video images, and further perform face recognition on the detected face images, compare the face database, and find the information of the participating personnel corresponding to the face. The feature recognition server can also deploy audio detection and voiceprint feature recognition algorithms to perform audio detection on the received audio, and further perform voiceprint feature recognition on the detected audio. Each of the plurality of participating terminals can be a computer, a mobile phone, a tablet computer, a conference all-in-one machine, etc., which is not limited herein. Each participating terminal can be a device integrating audio-video devices such as a camera and a microphone, or a device that can be connected to audio-video devices such as a camera and a microphone. In this way, each participating terminal can obtain corresponding audio-video stream data.

[0124] In a remote video conference, when there are multiple participating terminals communicatively connected, the server can receive audio-video stream data from each participating terminal. For example, when there are three participating terminals communicatively connected, the server can receive audio-video stream data from each of these three participating terminals. In this case, the server can receive three audio-video stream data of these three participating terminals. For another example, when there are five participating terminals communicatively connected, the server can receive audio-video stream data from each of these five participating terminals. In this case, the server can receive five audio-video stream data of these five participating terminals. Of course, the number of the plurality of participating terminals can be set according to actual applications, which is not limited herein.

[0125] S102: Determine the target audio-video stream data corresponding to at least one speaker from the at least one audio-video stream data;

[0126] In the specific implementation process, the at least one speaker may be the person who is speaking. As long as there is sound in the audio-video stream data of a certain participating terminal, it means that there is a speaker among the participants using that participating terminal. When the at least one speaker is one, it is possible to determine that the participant using the participating terminal is the current speaker who is speaking by detecting the audio information of the participating terminal. The specific detection is the same as the prior art and will not be limited here. In addition, the at least one speaker may also be multiple. In practical applications, the at least one speaker may be the host, or any other participant except the host, which will not be limited here.

[0127] In the specific implementation process, the target audio-video stream data may be one or multiple. When the target video stream data is multiple, correspondingly, the at least one speaker is multiple. Multiple speakers may be in different target audio-video stream data. For example, if the target audio-video stream data is three, the at least one speaker is three, and these three speakers may be respectively in these three different audio-video stream data. Another example is that if the target audio-video stream data is three, the at least one speaker is five, where three speakers are simultaneously in one target audio-video stream data, and the other two speakers are respectively in the other two different target audio-video stream data. In addition, when the target audio-video stream data is one and the at least one speaker is multiple, multiple speakers are in the same audio-video stream data. Of course, in practical applications, the relationship between the at least one speaker and the target audio-video stream data may also be other situations, which will not be elaborated here.

[0128] S103: Obtain the biometric information for identifying each of the at least one speaker from the target audio-video stream data, where the biometric information includes at least one of face feature information and voiceprint feature information;

[0129] In the specific implementation process, it is possible to obtain the face feature information of each of the at least one speaker from the target audio-video stream data by using a face recognition method, or to obtain the voiceprint feature information of each of the at least one speaker from the target audio-video stream data by using a voiceprint recognition method. The specific implementation of the face recognition method and the voiceprint recognition method is the same as the prior art and will not be elaborated here.

[0130] S104: Generate virtual business cards for each of the at least one speaker according to the biometric information of each of the at least one speaker;

[0131] In the specific implementation process, after obtaining the biometric information for identifying each of the at least one speaker from the target audio-video stream data, the virtual business cards for each of the at least one speaker can be generated according to the biometric information of each of the at least one speaker. Since the biometric information corresponding to different participants is different, the specific content of the generated virtual business cards is also different accordingly. For example, for speaker A, their business card corresponds to virtual business card a, and for speaker B, their business card corresponds to virtual business card b. In addition, if the biometric information of multiple different speakers is obtained from the target audio-video stream data, that is, the at least one speaker is multiple, accordingly, virtual business cards corresponding to their respective speakers will be generated according to the biometric information of the multiple different speakers. For example, when the at least one speaker is three, three virtual business cards will be generated, where each virtual business card corresponds to the speaker associated with the corresponding biometric information. Of course, the situation of the generated virtual business cards can also be other situations, which are not limited here.

[0132] S105: Overlay the virtual business cards of each of the at least one speaker onto the target audio-video stream data, and synthesize them with the other audio-video stream data in the multiple audio-video stream data except the target audio-video stream data into a single audio-video stream data to be sent;

[0133] In the specific implementation process, after generating the virtual business cards corresponding to the biometric information of each of the at least one speaker, overlay the virtual business cards of each of the at least one speaker onto the target audio-video stream data, and synthesize them with the other audio-video stream data in the multiple audio-video stream data except the target audio-video stream data into a single audio-video stream data to be sent. In this way, the virtual business cards of each of the at least one speaker can be shared with all participating terminals, thus ensuring the sharing efficiency of the business cards.

[0134] S106: Send the audio-video stream data to be sent to the multiple participating terminals so that the multiple participating terminals can display the virtual business cards of each of the at least one speaker.

[0135] In the specific implementation process, after synthesizing the target audio-video stream data with the virtual business cards of each of the at least one speaker and the other audio-video stream data into a single audio-video stream data to be sent, the synthesized audio-video stream data to be sent can be sent to the multiple participating terminals. In this way, each participating terminal among the multiple participating terminals can display the virtual business cards of each of the at least one speaker. In this way, the users using the corresponding participating terminals can share the business cards of all speakers with each other through the corresponding participating terminals, thus ensuring the business card sharing efficiency and improving the conference communication efficiency.

[0136] In an embodiment of the present disclosure, if multiple different voiceprint feature information are simultaneously obtained from the target audio-video stream data, the method further includes:

[0137] Determine that the at least one speaker is multiple speakers corresponding to the multiple different voiceprint feature information.

[0138] In a specific implementation process, if multiple different voiceprint feature information are simultaneously obtained from the target audio-video stream data, where different voiceprint feature information identifies different participants, at this time, the participants corresponding to the multiple different voiceprint feature information are speaking, that is, currently multiple people are speaking and there are multiple speakers speaking. Correspondingly, the at least one speaker is multiple speakers corresponding to the multiple different voiceprint feature information, and the multiple speakers can be speakers using the same participant terminal or speakers using different participant terminals. For example, three different voiceprint feature information S1, S2, and S3 are simultaneously obtained from the target audio-video stream data, where the voiceprint feature information S1 comes from the audio-video stream data of participant terminal 1, the voiceprint feature information S2 comes from the audio-video stream data of participant terminal 2, and the voiceprint feature information S3 comes from the audio-video stream data of participant terminal 3. Correspondingly, the speakers corresponding to the three different voiceprint feature information S1, S2, and S3 are speakers using different participant terminals. Of course, in practical applications, the situation where the at least one speaker is multiple can also be other situations, which will not be elaborated here. When it is recognized that the at least one speaker is multiple, subsequent business card sharing can display the business cards of multiple speakers, thereby improving the sharing efficiency.

[0139] In an embodiment of the present disclosure, if multiple different voiceprint feature information are successively obtained from the target audio-video stream data within a preset duration, the method further includes:

[0140] Determine that the at least one speaker is multiple speakers corresponding to the multiple different feature information.

[0141] In a specific implementation process, if multiple different voiceprint feature information are successively obtained from the target audio-video stream data within a preset duration, where the preset duration can be a duration set according to the actual usage habits of the user or a duration manually set by the user, which is not limited here. For example, the preset duration is 30s, and within 30s, three different voiceprint feature information S4, S5, and S6 are successively obtained from the target audio-video stream data. For example, it is a scenario of multiple people talking in the same participant terminal, or a scenario of multiple people talking in different participant terminals. At this time, there are multiple speakers. Subsequent business card sharing can display the business cards of multiple speakers, thereby improving the sharing efficiency.

[0142] In the embodiments of the present disclosure, as Figure 2 shown, step S105: superimposing the respective virtual business cards of the at least one speaker onto the target audio-visual stream data, includes:

[0143] S201: For each of the at least one speaker, detecting the coordinate position of the face region of the corresponding speaker from the target audio-visual stream data;

[0144] S202: Determining the target position and size of the virtual business card of the speaker according to the coordinate position of the face region of the speaker;

[0145] S203: Superimposing the virtual business card onto the target audio-visual stream data according to the target position and size of the virtual business card.

[0146] In the specific implementation process, the specific implementation processes of steps S201 to S203 are as follows:

[0147] First, when there is a face region in the target audio-visual stream data, for each of the at least one speaker, detecting the coordinate position of the face region of the corresponding speaker from the target audio-visual stream data, and then, determining the target position and size of the virtual business card of the speaker according to the coordinate position of the face region of the speaker. For example, according to the coordinate position of the face region of the speaker, calculating the position in front of the chest or the position on the top of the head of the speaker, and taking the calculated position in front of the chest as the target position of the virtual business card, or taking the calculated position on the top of the head as the target position of the virtual business card. For example, if the coordinate position of the face region of the speaker is (x0, y0), (x1, y0), (x0, y1), (x1, y1), the position 5 coordinate positions below the face region can be taken as the target position to superimpose the virtual business card, and the virtual business card can be superimposed starting from the coordinate position (x0, y1 + 5). In addition, when the size of the virtual business card is fixed, if the bottom of the virtual business card exceeds the image region of the corresponding video frame. For example, the height of the video frame is y and the height of the virtual business card is h. If after superimposing the virtual business card, y1 + 5 + h > y, the bottom of the virtual business card exceeds the video frame, then the content of the virtual business card will not be completely displayed. The coordinate position for superimposing the virtual business card can be adjusted from (x0, y1 + 5) to (x0, y - h), so that the bottom of the virtual business card is flush with the bottom of the corresponding video frame, thus ensuring the complete display of the virtual business card and ensuring the sharing quality of the virtual business card.

[0148] Then, according to the target position of the virtual business card, the virtual business card is superimposed on the target audio-visual stream data. For example, the virtual business card is superimposed on the chest position of the corresponding speaker. In the specific implementation process, the identity information of the corresponding speaker can be drawn on a semi-transparent business card picture to generate a virtual business card, and then, according to the target position, the virtual business card is superimposed on the target audio-visual stream data. Since the target position of the virtual business card is the position determined according to the coordinate position of the face area of the speaker, the virtual business card can be displayed at a suitable position of the speaker, thus ensuring the correct association between the speaker and the virtual business card and improving the sharing efficiency of the virtual business card.

[0149] In the embodiment of the present disclosure, if the face area of at least one speaker is not detected in the target audio-visual stream data, the method further includes:

[0150] Superimposing the respective virtual business cards of the at least one speaker on the target audio-visual stream data according to a preset coordinate position.

[0151] In the specific implementation process, the preset coordinate position can be a position preset by those skilled in the art according to actual application needs. If at least one speaker does not turn on the camera of the corresponding terminal, or at least one speaker turns his back to the camera of the corresponding terminal, or at least one speaker faces the camera of the corresponding terminal sideways, the association relationship between at least one speaker and the voiceprint feature information can be pre-recorded. When the face area of at least one speaker is not detected in the target audio-visual stream data, the corresponding virtual business card of each of the at least one speaker can be determined according to the voiceprint feature information of each of the at least one speaker, and then, the respective virtual business cards of the at least one speaker are superimposed on the target video stream data according to the preset coordinate position. For example, when the preset coordinate position is the lower right corner of the corresponding video screen, the virtual business card is displayed at the lower right corner of the corresponding video screen of the corresponding speaker. For another example, when the preset coordinate position is the lower left corner of the corresponding video screen, the virtual business card is displayed at the lower left corner of the corresponding video screen of the corresponding speaker. Of course, the specific position of the preset coordinate position can also be set according to actual application needs, which is not limited herein.

[0152] In the embodiment of the present disclosure, as Figure 3 shown, step S105: superimposing the respective virtual business cards of the at least one speaker on the target audio-visual stream data includes:

[0153] S301: Calculate the average gray level value corresponding to each color channel in the preset color channel of at least one image in the target audio-visual stream data, and adjust the chroma of the virtual business cards of each of the at least one speaker according to the ratio of the average gray level values corresponding to each color channel, so as to obtain the adjusted virtual business cards of each of the at least one speaker, so that the contrast between the chroma of the adjusted virtual business cards of each of the at least one speaker and the chroma of the at least one image is greater than a preset value;

[0154] S302: Superimpose the adjusted virtual business cards of each of the at least one speaker on the target audio-visual stream data.

[0155] In the specific implementation process, the specific implementation processes of steps S301 to S302 are as follows:

[0156] First, calculate the average gray level value corresponding to each color channel in the preset color channel of at least one image in the target audio-visual stream data. Among them, the at least one image can be one or multiple. The target audio-visual stream data often includes multiple images, and the at least one image can be selected from the multiple images included in the target audio-visual stream data. In addition, the preset color channel can be the RGB channel including the three color channels of red, green, and blue, and the preset color channel can also be the HSV channel including hue (H), saturation (S), and value (V). Of course, the preset color channel can also be preset according to actual application needs, which is not limited here. After setting the preset color channel, the average gray level value corresponding to each color channel included in the preset color channel of the at least one image can be calculated. Taking the at least one image as the target image and the preset color channel as the RGB channel as an example, the specific calculation process of the average gray level value corresponding to the target image in the R channel is to add the brightness of each pixel point on the target image in the R channel and then divide by the total number of pixel points of the target image. In this way, the average gray level value corresponding to the target image in the R channel is determined. Based on the same calculation principle, the average gray level value corresponding to the target image in the G channel and the average gray level value corresponding to the target image in the B channel can be calculated. In this way, the ratio of the average gray level values corresponding to each color channel is determined.

[0157] After determining the ratio of the average gray scale values corresponding to each color channel, the key color channel can be determined according to the average gray scale values corresponding to each color channel. Then, according to the key color channel, the chroma of the virtual business cards of each of the at least one speaker is adjusted so that the contrast between the chroma of the adjusted virtual business cards of each of the at least one speaker and the chroma of the at least one image is greater than a preset value. The preset value is a value preset according to actual application needs. For example, the preset value is 90%. For example, if the background of the at least one image is black and black fonts are still used to display the content in the virtual business card, the contrast between the virtual business card and the at least one image is small, and it is impossible for the participants to clearly determine the relevant content of the virtual business card, resulting in poor sharing quality of the virtual business card. Still taking the above RGB color channels as an example, if the average gray scale value corresponding to the R channel in the target image is greater than both the average gray scale value corresponding to the G channel and the average gray scale value corresponding to the B channel, the color channel with the opposite average gray scale value ratio can be used to adjust the chroma of the virtual business cards of each of the at least one speaker. For example, the gray scale distribution of the virtual business cards of each of the at least one speaker is adjusted, the ratio of the average gray scale value of the R channel of the virtual business cards of each of the at least one speaker is reduced, and the ratio of the average gray scale value of the G channel and the ratio of the average gray scale value of the B channel are correspondingly increased, thereby realizing the adjustment of the chroma of the virtual business cards of each of the at least one speaker.

[0158] After adjusting the chroma of the virtual business cards of each of the at least one speaker, the adjusted virtual business cards of each of the at least one speaker are obtained so that the contrast between the chroma of the adjusted virtual business cards of each of the at least one speaker and the chroma of the at least one image is greater than a preset value. Then, the adjusted virtual business cards of each of the at least one speaker are superimposed on the target audio-video stream data, thereby improving the contrast between the virtual business cards of each of the at least one speaker and the at least one image. For example, when the background of the at least one image is black, white fonts can be used to display the content in the virtual business card, so as to ensure the sharing effect of the virtual business card.

[0159] In addition, in the specific implementation process, multiple different formats of virtual business cards can be preset. For example, virtual business cards with different font sizes and virtual business cards with different font colors. In the specific implementation process, according to the ratio of the average gray scale values corresponding to each color channel in the preset color channel of at least one image in the target audio-video stream data, a virtual business card with better contrast can be selected from multiple different formats of virtual business cards and superimposed on the target audio-video stream data, thereby ensuring the sharing effect of the virtual business card.

[0160] In the embodiments of the present disclosure, such as Figure 4As shown in the figure, step S202: Determine the target position and size of the virtual business card of the speaker according to the coordinate position of the face area of the speaker, including:

[0161] S401: Determine the coordinate positions of other areas in the target audio-video stream data except the face area of the speaker according to the coordinate position of the face area of the speaker;

[0162] S402: Determine at least one connected area from the other areas according to the coordinate positions of the other areas;

[0163] S403: Determine a target connected area larger than a preset area from the at least one connected area, and the coordinate position of the target connected area;

[0164] S404: Determine the target position and size of the virtual business card of the speaker based on the target connected area and the coordinate position of the target connected area.

[0165] In the specific implementation process, the specific implementation processes of steps S401 to S404 are as follows:

[0166] First, determine the coordinate positions of other areas in the target audio-video stream data except the face area of the speaker according to the coordinate position of the face area of the speaker. For example, in addition to the face area of the speaker, area C is also included. Then, determine at least one connected area from the other areas according to the coordinate positions of the other areas. For example, area C includes four connected domains: area c1, area c2, area c3, and area c4. Then, determine a target connected domain larger than a preset area from the at least one connected domain, and the coordinate position of the target connected domain. The target connected domain can be the area with the largest area among the at least one connected domain. For example, determine the connected domain with the largest area among areas c1 to c4 as c1, and the coordinate position of this connected domain c1. Then, the connected domain c1 can be used as the target connected domain. Then, determine the target position and size of the virtual business card of the speaker according to the coordinate position of the target connected domain. In this way, the area with the largest remaining connected domain in the target audio-video stream data can be selected to overlay the virtual business card to ensure the complete display of the virtual business card, thereby ensuring the display quality of the virtual business card.

[0167] In the embodiment of the present disclosure, step S404: Determine the target position and size of the virtual business card of the speaker based on the target connected area and the coordinate position of the target connected area. There can be the following two implementation manners, but are not limited to the following two implementation manners. The first implementation manner includes:

[0168] According to the preset business card shape, determine the region with the largest area in the target connected region that is the same as the preset business card shape, and adjust the target position and size of the virtual business card of the speaker according to the coordinate position where the region with the largest area is located.

[0169] In the specific implementation process, the preset business card shape can be the default business card shape of the system, or the business card shape manually set by the conference administrator. The preset business card shape can be one of a right-angled rectangle, a rounded rectangle, a circle, a triangle, a trapezoid, and a square, which is not limited here. When the preset business card shape is determined, the region with the largest area that is the same as the preset business card shape can be determined from the target connected region. For example, when the preset business card is a rounded rectangle, the region with the largest area that is the same as the rounded rectangle shape is determined from the target connected region. The shape of the region with the largest area is also a rounded rectangle. At this time, the target position and size of the virtual business card of the speaker can be adjusted according to the coordinate position where the region with the largest area is located. It can be that the virtual business card of the speaker fills the region with the largest area, or the virtual business card of the speaker is set at the central region position of the region with the largest area, and there is a certain distance between the region occupied by the virtual business card of the speaker and the edge of the region with the largest area. Thus, the target position and size of the virtual business card of the speaker are adjusted according to the region with the largest area, ensuring better sharing quality of the virtual business card of the speaker.

[0170] In the embodiment of the present disclosure, the second implementation manner of step S404 is as Figure 5 shown. Specifically, step S404: Determine the target position and size of the virtual business card of the speaker based on the target connected region and the coordinate position of the target connected region, including:

[0171] S501: Determine the inscribed figure with the largest area within the target connected region;

[0172] S502: Use the inscribed figure with the largest area as the shape of the virtual business card of the speaker;

[0173] S503: Adjust the target position and size of the virtual business card of the speaker according to the coordinate position where the inscribed figure with the largest area is located.

[0174] In the specific implementation process, the specific implementation processes of steps S501 to S503 are as follows:

[0175] First, determine the inscribed figure with the largest area within the target connected region. It can be to determine the inscribed figure with the largest area in the shape of a business card from within the target connected region according to a pre-set business card shape, or directly determine the inscribed figure with the largest area from within the target connected region. The pre-set business card shape can be that the system pre-sets multiple business card shapes, or various business card shapes manually set by the conference administrator. For example, the multiple business card shapes include at least two of a right-angled rectangle, a rounded rectangle, a circle, an ellipse, a triangle, a trapezoid, and a square. Of course, it can also be a combination of other multiple shapes, which is not limited here. After determining the inscribed figure with the largest area within the target connected region, use the coordinate position where the inscribed figure with the largest area is located as the shape of the virtual business card of the speaker, and adjust the target position and size of the virtual business card of the speaker according to the coordinate position where the inscribed figure with the largest area is located. It can be to fill the area where the inscribed figure with the largest area is located with the virtual business card of the speaker, or set the virtual business card of the speaker in the central area of the area where the inscribed figure with the largest area is located, and there is a certain distance between the area occupied by the virtual business card of the speaker and the edge of the area where the inscribed figure with the largest area is located. Thus, the target position and size of the virtual business card of the speaker are adjusted according to the inscribed figure with the largest area, and further, when the target connected region is fixed, the display of the virtual business card is maximized, and further, the display quality of the virtual business card is ensured.

[0176] In the embodiment of the present disclosure, in step S105: superimposing the respective virtual business cards of the at least one speaker on the target audio-video stream data and synthesizing the target audio-video stream data with other audio-video stream data in the multiple audio-video stream data into a to-be-sent audio-video stream data includes:

[0177] According to a preset layout, synthesize the target audio-video stream data with the respective virtual business cards of the at least one speaker superimposed thereon and other audio-video stream data in the multiple audio-video stream data except the target audio-video stream data into a to-be-sent audio-video stream data, so that the video picture corresponding to the target audio-video stream data with the respective virtual business cards of the at least one speaker superimposed thereon among the multiple participating terminals is larger than the video pictures corresponding to the other audio-video stream data.

[0178] In the specific implementation process, according to the preset layout, the target audio-visual stream data with the virtual business cards of each of the at least one speaker superimposed thereon can be synthesized with the other audio-visual stream data in the multiple audio-visual stream data except the target audio-visual stream data into a to-be-transmitted audio-visual stream data. It can be to superimpose the virtual business cards of each of the at least one speaker on the target audio-visual stream data, and then synthesize the target audio-visual stream data with the virtual business cards of each of the at least one speaker superimposed thereon with the other audio-visual stream data in the multiple audio-visual stream data except the target audio-visual stream data into a to-be-transmitted audio-visual stream data. In the specific implementation process, the preset layout can be that the video screens of the terminals corresponding to the at least one speaker are displayed in a large screen, and the participating terminals corresponding to the other participants except the at least one speaker are displayed in small screens. As Figure 6 Shown is one of the schematic diagrams of the preset layout when the at least one speaker is one. Among them, the video screen of the speaker F is shown in a large view, and the video screens of the other participants R1 to R7 using different participating terminals are shown at the bottom and on the right around the video screen of the speaker F. In this way, after each participating terminal in the multiple participating terminals receives the to-be-transmitted audio-visual stream data, the video screen corresponding to the target audio-visual stream data with the virtual business card superimposed thereon in each participating terminal will be displayed larger than the video screen corresponding to the other audio-visual stream data. As Figure 7 Shown is that the at least one speaker is two, and these two speakers, such as P1 and P2, appear in the video screens of two different participating terminals respectively. The view sizes of the video screens corresponding to the speaker P1 and the video screen corresponding to the speaker P2 are the same, and the video screens of the other participants P3 to P8 are shown at the bottom of the video screens of the speaker P1 and P3. Of course, the video screens of the at least one speaker can also be set according to actual application needs, which will not be elaborated here.

[0179] In the embodiment of the present disclosure, before step S101: receiving multiple audio-visual stream data of multiple participating terminals in communication connection, the method further includes:

[0180] Pre-storing the corresponding relationship between the biometric information of known participants and the virtual business cards.

[0181] In the specific implementation process, before receiving the audio and video stream data of multiple participating terminals in a communication connection, the corresponding relationship between the biometric information of known participants and virtual business cards is pre-stored. It can be the pre-entry of the corresponding relationship between the face pictures of known participants and virtual business cards, and / or the pre-entry of the corresponding relationship between the audio files of known participants and virtual business cards, and this corresponding relationship is stored. For example, before entering the remote video conferencing system, the face photos of each participant and the corresponding personnel information (virtual business cards) are entered into the system. It can be that the background administrator logs in to the remote video conferencing system and submits the face photos of the participants and the corresponding personnel information (virtual business cards), or when each participant logs in to the remote video conferencing system by themselves, they mention their own face photos and personnel information (virtual business cards). Another example is that before entering the remote video conferencing system, the audio files of each participant and the corresponding personnel information (virtual business cards) are entered into the system. In this way, when the speaker is a known participant with a known face photo and virtual business card, or when the speaker is a known participant with a known audio file and virtual business card, the virtual business card of the speaker can be quickly determined according to the pre-stored corresponding relationship between the biometric information of known participants and virtual business cards, thus ensuring the sharing efficiency of virtual business cards.

[0182] In the embodiments of the present disclosure, as Figure 8 shown, if the biometric information of the at least one speaker is not obtained in the target audio and video stream data, the method further includes:

[0183] S601: Determine the respective video images of the at least one speaker;

[0184] S602: Receive the screenshot operation of the respective video images of the at least one speaker by the person with entry permission, and in response to the screenshot operation, determine the biometric information of the at least one speaker;

[0185] S603: Receive the text input operation for the respective video images of the at least one speaker, and in response to the text input operation, determine the virtual business cards of the at least one speaker;

[0186] S604: Associate the biometric information of the at least one speaker with the virtual business cards of the at least one speaker.

[0187] In the specific implementation process, the specific implementation processes of steps S601 to S604 are as follows:

[0188] If the multiple audio-video stream data do not obtain the biometric information of the at least one speaker, for example, if all of the at least one speakers are those who temporarily join the remote video conferencing system and their face pictures and virtual business cards are not pre-entered into the system, the video frames of the at least one speaker can be determined first. For example, if the biometric information of the at least one speaker is not recognized in the multiple audio-video stream data, "unknown person" can be marked in the video frames of the at least one speaker respectively to determine the video frames of the at least one speaker. Then, receive the screenshot operation of the person with the input permission for the video frames of the at least one speaker respectively, and in response to the screenshot operation, determine the virtual business cards of the at least one speaker respectively. The person with the input permission can be the meeting host or the meeting organizer, which is not limited herein. The screenshot operation can be the operation of the person with the input permission to circle the face area of the at least one speaker in the video frame. For example, circle the face area of the at least one speaker with a circle, or click the video frame corresponding to the at least one speaker with the mouse. Of course, those skilled in the art can also set the specific form of the screenshot operation according to the actual application needs, which is not limited herein.

[0189] After the person with the input permission performs the screenshot operation on the video frames of the at least one speaker respectively, receive the text input operation on the video frames of the at least one speaker respectively, and in response to the text input operation, determine the virtual business cards of the at least one speaker respectively. It can be that after the person with the input permission performs the screenshot operation on the video frames of the at least one speaker respectively, a text input box for entering the virtual business cards of the at least one speaker respectively pops up. The person with the input permission can enter the virtual business cards of the at least one speaker respectively in this text input box. For example, enter the names, positions, departments, contact information, etc. of the at least one speaker respectively. Then, associate the biometric information of the at least one speaker respectively with the virtual business cards of the at least one speaker respectively. In this way, the biometric information and virtual business cards of the speakers who temporarily join the video conference can be entered in real time, so as to ensure the sharing of the virtual business cards of any participant, and further improve the sharing efficiency of the virtual business cards.

[0190] In the embodiment of the present disclosure, if no speaker is detected in the at least one audio-video stream data, the method further includes:

[0191] Determine the conference terminal corresponding to the meeting host from the multiple participating terminals, and use the audio-video stream data corresponding to the conference terminal as the target audio-video stream data.

[0192] In the specific implementation process, if no speaker is detected in the at least one audio-video stream data, that is to say, none of the participants corresponding to the current participating terminals are speaking, that is, there is no speaker, the meeting terminal corresponding to the meeting host can be determined from the multiple participating terminals, and the audio-video stream data corresponding to the meeting terminal can be used as the target audio-video stream data. In this way, if there is no speaker currently, the virtual business card of the meeting host can be displayed on the terminal corresponding to the meeting host. In this case, all participants can know the virtual business card of the meeting host, ensuring the sharing efficiency of the virtual business card. Of course, the default display situation of the virtual business card when there is no speaker can also be set according to actual application needs, which will not be elaborated here.

[0193] Based on the same general concept, as Figure 9 shown, an embodiment of the present disclosure further provides a virtual business card sending device, which is applied to a server. The device includes:

[0194] A receiving unit 10, configured to receive at least one audio-video stream data of multiple participating terminals communicatively connected to the server;

[0195] A determining unit 20, configured to determine target audio-video stream data corresponding to at least one speaker from the at least one audio-video stream data;

[0196] An obtaining unit 30, configured to obtain biometric information for identifying each of the at least one speaker from the target audio-video stream data, where the biometric information includes at least one of face feature information and voiceprint feature information;

[0197] A generating unit 40, configured to generate virtual business cards for each of the at least one speaker according to the biometric information of each of the at least one speaker;

[0198] A synthesizing unit 50, configured to superimpose the virtual business cards of each of the at least one speaker onto the target audio-video stream data, and synthesize the target audio-video stream data with other audio-video stream data except the target audio-video stream data in the multiple audio-video stream data into a to-be-sent audio-video stream data;

[0199] A sending unit 60, configured to send the to-be-sent audio-video stream data to the multiple participating terminals, so that the multiple participating terminals display the virtual business cards of each of the at least one speaker.

[0200] In the embodiment of the present disclosure, if the obtaining unit 30 simultaneously obtains multiple different voiceprint feature information from the target audio-video stream data, the determining unit 20 is further configured to:

[0201] Determine that the at least one speaker is multiple speakers corresponding to the multiple different voiceprint feature information.

[0202] In an embodiment of the present disclosure, if the obtaining unit 30 successively obtains a plurality of different voiceprint feature information from the target audio-visual stream data, the determining unit 20 is further configured to:

[0203] Determine that the at least one speaker is a plurality of speakers corresponding to the plurality of different feature information.

[0204] In an embodiment of the present disclosure, the synthesizing unit 50 is configured to:

[0205] For each of the at least one speaker, detect the coordinate position of the face region of the corresponding speaker from the target audio-visual stream data;

[0206] Determine the target position and size of the virtual business card of the speaker according to the coordinate position of the face region of the speaker;

[0207] Overlay the virtual business card to the target audio-visual stream data according to the target position and size of the virtual business card. In an embodiment of the present disclosure, if the face region of the speaker is not detected in the target audio-visual stream data, the synthesizing unit 50 is further configured to:

[0208] Overlay the virtual business cards of the at least one speaker to the target audio-visual stream data according to a preset coordinate position.

[0209] In an embodiment of the present disclosure, the synthesizing unit 50 is configured to:

[0210] Calculate the average gray scale value corresponding to each color channel in a preset color channel of at least one image in the target audio-visual stream data, and adjust the chroma of the virtual business cards of the at least one speaker according to the ratio of the average gray scale values corresponding to the respective color channels, so as to obtain the adjusted virtual business cards of the at least one speaker, so that the contrast between the chroma of the adjusted virtual business cards of the at least one speaker and the chroma of the at least one image is greater than a preset value;

[0211] Overlay the adjusted virtual business cards of the at least one speaker to the target audio-visual stream data.

[0212] In an embodiment of the present disclosure, the synthesizing unit 50 is configured to:

[0213] Determine the coordinate position of other regions in the target audio-visual stream data except the face region of the speaker according to the coordinate position of the face region of the speaker;

[0214] Determine at least one connected region from the other regions according to the coordinate position of the other regions;

[0215] Determine a target connected region with an area greater than a preset area from the at least one connected region, and the coordinate position of the target connected region;

[0216] Determine the target position and size of the virtual business card of the speaker according to the target connected region and the position of the target connected region.

[0217] In an embodiment of the present disclosure, the synthesis unit 50 is configured to:

[0218] Determine the region with the largest area in the target connected region that is the same as the preset business card shape according to the preset business card shape, and adjust the target position and size of the virtual business card of the speaker according to the coordinate position where the region with the largest area is located.

[0219] In an embodiment of the present disclosure, the synthesis unit 50 is configured to:

[0220] Determine the inscribed figure with the largest area within the target connected region;

[0221] Take the inscribed figure with the largest area as the shape of the virtual business card of the speaker;

[0222] Adjust the target position and size of the virtual business card of the speaker according to the coordinate position where the inscribed figure with the largest area is located.

[0223] In an embodiment of the present disclosure, the synthesis unit 50 is configured to:

[0224] According to a preset layout, synthesize the target audio-visual stream data with the virtual business cards of the at least one speaker superimposed thereon and the other audio-visual stream data in the plurality of audio-visual stream data except the target audio-visual stream data into a to-be-transmitted audio-visual stream data, so that the video picture corresponding to the target audio-visual stream data with the virtual business cards of the at least one speaker superimposed thereon in the plurality of participating terminals is larger than the video picture corresponding to the other audio-visual stream data.

[0225] In an embodiment of the present disclosure, the device further includes a storage unit, and the storage unit is configured to:

[0226] Pre-store the correspondence between the biometric information of known participants and virtual business cards.

[0227] In an embodiment of the present disclosure, if the biometric information of the at least one speaker is not obtained in the target audio-visual stream data, the device further includes an input unit, and the input unit is configured to:

[0228] Determine the video pictures of the at least one speaker respectively;

[0229] Receive the screenshot operation of the video images of the at least one speaker by the personnel with entry permission, and in response to the screenshot operation, determine the biometric information of the at least one speaker respectively;

[0230] Receive the text input operation for the video images of the at least one speaker respectively, and in response to the text input operation, determine the virtual business cards of the at least one speaker respectively;

[0231] Associate the biometric information of the at least one speaker respectively with the virtual business cards of the at least one speaker respectively.

[0232] In the embodiment of the present disclosure, if no speaker is detected in the multiple audio-video stream data, the device further includes a setting unit, and the setting unit is configured to:

[0233] Determine the conference terminal corresponding to the conference host from the multiple participating terminals, and use the audio-video stream data corresponding to the conference terminal as the target audio-video stream data. Based on the same general inventive concept, as Figure 10 shown, the embodiment of the present disclosure further provides a virtual business card sending system, wherein the sending system includes a server 70 and multiple participating terminals 80 communicatively connected to the server 80;

[0234] The server 70 is configured to receive at least one audio-video stream data of multiple participating terminals communicatively connected to the server;

[0235] The server 70 is further configured to determine the target audio-video stream data corresponding to at least one speaker from the at least one audio-video stream data;

[0236] The server 70 is further configured to obtain from the target audio-video stream data the biometric information for identifying the at least one speaker respectively, where the biometric information includes at least one of face feature information and voiceprint feature information;

[0237] The server 70 is further configured to generate the virtual business cards of the at least one speaker respectively according to the biometric information of the at least one speaker respectively; superimpose the virtual business cards of the at least one speaker respectively on the target audio-video stream data, and synthesize the target audio-video stream data with other audio-video stream data in the multiple audio-video stream data into a to-be-sent audio-video stream data; send the to-be-sent audio-video stream data to the multiple participating terminals 80;

[0238] The multiple participating terminals 80 are configured to display the virtual business cards of the at least one speaker respectively.

[0239] In the implementation of the present disclosure, the server 70 includes a conference server 701 and a feature recognition server 702 communicatively connected to the conference server 701. Among them, the conference server 701 is configured to receive the plurality of audio-visual stream data, determine the target audio-visual stream data from the plurality of audio-visual stream data, and determine virtual business cards corresponding to the biometric information of each of the at least one speaker by invoking the feature detection and recognition interface of the feature recognition server, and send the to-be-transmitted audio-visual stream data with the virtual business cards superimposed thereon to the plurality of participating terminals 80;

[0240] The feature recognition server 702 is configured to identify the biometric information for identifying each of the at least one speaker from the target audio-visual stream data, and send the biometric information of each of the at least one speaker to the conference server 701.

[0241] In the embodiment of the present disclosure, the plurality of participating terminals 80 are further configured to:

[0242] Display the video picture corresponding to the target audio-visual stream data with the virtual business card superimposed thereon in a larger size than the video pictures corresponding to the other audio-visual stream data except the target audio stream data among the plurality of audio-visual stream data.

[0243] Based on the same inventive concept, as Figure 11 shown, the embodiment of the present disclosure further provides a virtual business card sending device, which includes:

[0244] A memory 90 and a processor 100;

[0245] Wherein, the memory 90 is used to store a computer program;

[0246] The processor 100 is used to execute the computer program in the memory to implement the following steps:

[0247] Receive at least one audio-visual stream data of a plurality of participating terminals communicatively connected to the server;

[0248] Determine the target audio-visual stream data corresponding to at least one speaker from the at least one audio-visual stream data;

[0249] Obtain the biometric information for identifying each of the at least one speaker from the target audio-visual stream data, where the biometric information includes at least one of face feature information and voiceprint feature information;

[0250] Generate virtual business cards for each of the at least one speaker according to the biometric information of each of the at least one speaker;

[0251] Overlay the respective virtual business cards of the at least one speaker onto the target audio-video stream data, and synthesize the target audio-video stream data with other audio-video stream data among the multiple audio-video stream data into a to-be-sent audio-video stream data;

[0252] Send the to-be-sent audio-video stream data to the multiple participating terminals, so that the multiple participating terminals display the respective virtual business cards of the at least one speaker.

[0253] Based on the same general inventive concept, an embodiment of the present disclosure further provides a computer non-transitory readable storage medium, wherein:

[0254] The storage medium stores computer instructions, and when the computer instructions run on a computer, the computer is caused to execute the method for sending a virtual business card as described above.

[0255] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0256] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0257] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0258] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus, causing a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one flow Figure 1 one or more flows and / or boxes Figure 1 or boxes.

[0259] Although the preferred embodiments of the present disclosure have been described, additional changes and modifications to these embodiments can be made by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to cover the preferred embodiments as well as all changes and modifications falling within the scope of the present disclosure.

[0260] Obviously, those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure is also intended to include these modifications and variations.

Claims

1. A method for sending virtual business cards, applied to a server, wherein, it includes: Receiving at least one audio-video stream data of multiple participating terminals communicatively connected to the server; Determining, from the at least one audio-video stream data, target audio-video stream data corresponding to at least one speaker; Obtaining, from the target audio-video stream data, biometric information for identifying each of the at least one speaker, where the biometric information includes at least one of face feature information and voiceprint feature information; Generating respective virtual business cards for each of the at least one speaker according to the biometric information of each of the at least one speaker; Overlaying the respective virtual business cards of the at least one speaker onto the target audio-video stream data, and synthesizing the target audio-video stream data with other audio-video stream data other than the target audio-video stream data among the multiple audio-video stream data into a to-be-sent audio-video stream data; Sending the to-be-sent audio-video stream data to the multiple participating terminals, so that the multiple participating terminals display the respective virtual business cards of the at least one speaker; wherein, the overlaying the respective virtual business cards of the at least one speaker onto the target audio-video stream data includes: For each of the at least one speaker, detecting the coordinate position of the face area of the corresponding speaker from the target audio-video stream data; Determining the target position and size of the virtual business card of the speaker according to the coordinate position of the face area of the speaker; Overlaying the virtual business card onto the target audio-video stream data according to the target position and size of the virtual business card; wherein, the determining the target position and size of the virtual business card of the speaker according to the coordinate position of the face area of the speaker includes: Determining the coordinate position of other areas in the target audio-video stream data except the face area of the speaker according to the coordinate position of the face area of the speaker; Determining at least one connected area from the other areas according to the coordinate position of the other areas; Determining a target connected area larger than a preset area and the coordinate position of the target connected area from the at least one connected area; Determining the target position and size of the virtual business card of the speaker according to the target connected area and the coordinate position of the target connected area.

2. The method according to claim 1, wherein, if multiple different voiceprint feature information are simultaneously obtained from the target audio-video stream data, the method further includes: Determining that the at least one speaker is multiple speakers corresponding to the multiple different voiceprint feature information.

3. The method according to claim 1, wherein, if multiple different voiceprint feature information are successively obtained from the target audio-video stream data within a preset duration, the method further includes: Determining that the at least one speaker is multiple speakers corresponding to the multiple different feature information.

4. The method according to claim 1, wherein, if the face area of the at least one speaker is not detected in the target audio-video stream data, the method further includes: Superimpose the respective virtual business cards of the at least one speaker onto the target audio-video stream data according to the preset coordinate positions.

5. The method according to any one of claims 1-4, wherein, the superimposing the respective virtual business cards of the at least one speaker onto the target audio-video stream data includes: Calculating the average grayscale value corresponding to each color channel in a preset color channel of at least one image in the target audio-video stream data, and adjusting the chromaticity of the respective virtual business cards of the at least one speaker according to the ratio of the average grayscale values corresponding to the respective color channels, to obtain the respective adjusted virtual business cards of the at least one speaker, so that the contrast between the chromaticity of the respective adjusted virtual business cards of the at least one speaker and the chromaticity of the at least one image is greater than a preset value; Superimpose the respective adjusted virtual business cards of the at least one speaker onto the target audio-video stream data.

6. The method according to claim 1, wherein, the determining the target position and size of the virtual business card of the speaker according to the target connected region and the coordinate position of the target connected region includes: Determining the region with the largest area in the target connected region that is the same as the preset business card shape according to the preset business card shape, and adjusting the target position and size of the virtual business card of the speaker according to the coordinate position where the region with the largest area is located.

7. The method according to claim 1, wherein, the determining the target position and size of the virtual business card of the speaker according to the target connected region and the coordinate position of the target connected region includes: Determining the inscribed figure with the largest area within the target connected region; Taking the inscribed figure with the largest area as the shape of the virtual business card of the speaker; Adjusting the target position and size of the virtual business card of the speaker according to the coordinate position where the inscribed figure with the largest area is located.

8. The method according to any one of claims 1-4, wherein, the superimposing the respective virtual business cards of the at least one speaker onto the target audio-video stream data and synthesizing the target audio-video stream data with other audio-video stream data in the plurality of audio-video stream data except the target audio-video stream data into a single audio-video stream data to be sent includes: Synthesizing the target audio-video stream data with the respective virtual business cards of the at least one speaker superimposed thereon and other audio-video stream data in the plurality of audio-video stream data except the target audio-video stream data into a single audio-video stream data to be sent according to a preset layout, so that the video picture corresponding to the target audio-video stream data with the respective virtual business cards of the at least one speaker superimposed thereon among the plurality of participating terminals is larger than the video pictures corresponding to the other audio-video stream data.

9. The method according to claim 1, wherein, before receiving the plurality of audio-video stream data of the plurality of participating terminals in the communication connection, the method further includes: Pre-storing the corresponding relationship between the biometric information of known participants and the virtual business cards.

10. The method according to claim 1, wherein, if the biometric information of the at least one speaker is not obtained in the target audio-video stream data, the method further includes: Determine the respective video images of the at least one speaker; Receive a screenshot operation of the respective video images of the at least one speaker by a person with entry permission, and in response to the screenshot operation, determine the respective biometric information of the at least one speaker; Receive a text input operation for the respective video images of the at least one speaker, and in response to the text input operation, determine the respective virtual business cards of the at least one speaker; Associate the respective biometric information of the at least one speaker with the respective virtual business cards of the at least one speaker.

11. A virtual business card sending device, applied to a server, wherein, comprises: a receiving unit, configured to receive at least one audio-video stream data of a plurality of participating terminals communicatively connected to the server; a determining unit, configured to determine, from the plurality of audio-video stream data, target audio-video stream data corresponding to at least one speaker; an obtaining unit, configured to obtain, from the target audio-video stream data, biometric information for identifying the respective at least one speaker, where the biometric information includes at least one of face feature information and voiceprint feature information; a generating unit, configured to generate the respective virtual business cards of the at least one speaker according to the respective biometric information of the at least one speaker; a synthesizing unit, configured to superimpose the respective virtual business cards of the at least one speaker on the target audio-video stream data, and synthesize the target audio-video stream data with other audio-video stream data other than the target audio-video stream data in the plurality of audio-video stream data into a to-be-sent audio-video stream data; a sending unit, configured to send the to-be-sent audio-video stream data to the plurality of participating terminals, so that the plurality of participating terminals display the respective virtual business cards of the at least one speaker; wherein, the synthesizing unit is configured to: For each of the at least one speaker, detect the coordinate position of the face area of the corresponding speaker from the target audio-video stream data; Determine the size of the target position of the virtual business card of the speaker according to the coordinate position of the face area of the speaker; Superimpose the virtual business card on the target audio-video stream data according to the target position and size of the virtual business card; and, Determine the coordinate positions of other areas in the target audio-video stream data except the face area of the speaker according to the coordinate position of the face area of the speaker; Determine at least one connected area from the other areas according to the coordinate positions of the other areas; Determine a target connected area with an area larger than a preset area and the coordinate position of the target connected area from the at least one connected area; Determine the target position and size of the virtual business card of the speaker according to the target connected area and the position of the target connected area.

12. The device according to claim 11, wherein, If the obtaining unit simultaneously obtains a plurality of different voiceprint feature information from the target audio-video stream data, the determining unit is further configured to: Determine that the at least one speaker is a plurality of speakers corresponding to the plurality of different voiceprint feature information.

13. The device according to claim 11, wherein, If the obtaining unit successively obtains multiple different voiceprint feature information from the target audio-video stream data, the determining unit is further configured to: Determine that the at least one speaker is multiple speakers corresponding to the multiple different feature information.

14. A virtual business card sending system, Wherein, The sending system includes a server and multiple participating terminals communicatively connected to the server; The server is configured to receive at least one audio-video stream data of multiple participating terminals communicatively connected to the server; The server is further configured to determine, from the multiple audio-video stream data, target audio-video stream data corresponding to at least one speaker; The server is further configured to obtain biometric information for identifying each of the at least one speaker from the target audio-video stream data, where the biometric information includes at least one of face feature information and voiceprint feature information; The server is further configured to generate virtual business cards for each of the at least one speaker according to the biometric information of each of the at least one speaker; superimpose the virtual business cards of each of the at least one speaker on the target audio-video stream data, and synthesize the target audio-video stream data with other audio-video stream data other than the target audio-video stream data in the multiple audio-video stream data into a to-be-sent audio-video stream data; send the to-be-sent audio-video stream data to the multiple participating terminals; The multiple participating terminals are configured to display the virtual business cards of each of the at least one speaker; Wherein, the server is further configured to: For each of the at least one speaker, detect the coordinate position of the face area of the corresponding speaker from the target audio-video stream data; Determine the target position and size of the virtual business card of the speaker according to the coordinate position of the face area of the speaker; Superimpose the virtual business card on the target audio-video stream data according to the target position and size of the virtual business card; and, Determine the coordinate positions of other areas in the target audio-video stream data except the face area of the speaker according to the coordinate position of the face area of the speaker; Determine at least one connected area from the other areas according to the coordinate positions of the other areas; Determine a target connected area larger than a preset area and the coordinate position of the target connected area from the at least one connected area; Determine the target position and size of the virtual business card of the speaker according to the target connected area and the coordinate position of the target connected area.

15. The system according to claim 14, Wherein, The server includes a conference server and a feature recognition server communicatively connected to the conference server. The conference server is configured to receive the multiple audio-video stream data, determine the target audio-video stream data from the multiple audio-video stream data, and determine virtual business cards corresponding to the biometric information of each of the at least one speaker by calling the feature detection and recognition interface of the feature recognition server, and send the to-be-sent audio-video stream data with the virtual business card superimposed to the multiple participating terminals; The feature recognition server is configured to identify biometric information for identifying each of the at least one speaker from the target audio-visual stream data, and send the biometric information of each of the at least one speaker to the conference server.

16. The system according to claim 15, wherein, the plurality of participating terminals are further configured to: display the video picture corresponding to the target audio-visual stream data with the virtual business card superimposed thereon, at a size larger than the video pictures corresponding to the other audio-visual stream data in the plurality of audio-visual stream data except the target audio-visual stream data.

17. A virtual business card sending device, wherein, comprising: a memory and a processor; wherein, the memory is used to store a computer program; the processor is used to execute the computer program in the memory to implement the following steps: receive at least one audio-visual stream data of a plurality of participating terminals communicatively connected to a server; determine, from the plurality of audio-visual stream data, the target audio-visual stream data corresponding to at least one speaker; obtain biometric information for identifying each of the at least one speaker from the target audio-visual stream data, the biometric information including at least one of face feature information and voiceprint feature information; generate virtual business cards for each of the at least one speaker according to the biometric information of each of the at least one speaker; superimpose the virtual business cards of each of the at least one speaker on the target audio-visual stream data, and synthesize the target audio-visual stream data with the other audio-visual stream data in the plurality of audio-visual stream data except the target audio-visual stream data into a pending transmission audio-visual stream data; send the pending transmission audio-visual stream data to the plurality of participating terminals, so that the plurality of participating terminals display the virtual business cards of each of the at least one speaker; wherein, the superimposing the virtual business cards of each of the at least one speaker on the target audio-visual stream data includes: for each of the at least one speaker, detect the coordinate position of the face area of the corresponding speaker from the target audio-visual stream data; determine the target position and size of the virtual business card of the speaker according to the coordinate position of the face area of the speaker; superimpose the virtual business card on the target audio-visual stream data according to the target position and size of the virtual business card; wherein, the determining the target position and size of the virtual business card of the speaker according to the coordinate position of the face area of the speaker includes determine the coordinate position of other areas in the target audio-visual stream data except the face area of the speaker according to the coordinate position of the face area of the speaker; determine at least one connected area from the other areas according to the coordinate position of the other areas; determine a target connected area larger than a preset area and the coordinate position of the target connected area from the at least one connected area; determine the target position and size of the virtual business card of the speaker according to the target connected area and the coordinate position of the target connected area.

18. A computer non-transitory readable storage medium, wherein: The storage medium stores computer instructions, which, when run on a computer, cause the computer to execute the method for sending a virtual business card as described in any one of claims 1-10.

Citation Information

Patent Citations

  • Video conference data processing method and platform

    CN107333090A

  • Method and device for prompting information of participant in video conference

    CN107370981A

  • Method and device for pushing business card information based on video conference

    CN110519546A

  • Video conference method, system and device and storage medium

    CN110572607A