Video picture construction method and electronic device
Patent Information
- Application Number
- CN202210660465.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-01-24
- Filing Date
- 2022-06-13
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-06-13
AI Technical Summary
[0002]在目前视频会议的技术中,若要请与会人员发言,通常直接透过声音说出名字,但在多人会议中,因众多交谈声音,与会人员容易因为会议上的各种状况而忽略,进而导致会议中断
[0014]In summary, this disclosure document uses facial recognition to identify speakers with high display priority in video conferencing. Based on whether the multiple facial frames are speaking and the multiple display priorities, it determines that at least one of the multiple facial frames constitutes the main display area of the video screen, thereby enabling participants to more clearly see and notice the speaker's message.
Smart Images

Figure CN116524554B_ABST
Abstract
Description
Technical Field
[0001] This case relates to a method of composing video images, and more particularly to a method of composing video images and an electronic device. Background Technology
[0002] In current video conferencing technology, calling on a participant to speak usually involves calling their name. However, in multi-person conferences, due to numerous conversations, participants can easily miss certain words due to various circumstances, leading to conference interruptions. Furthermore, the speaker in a multi-person conference may also be overlooked. Additionally, video conferencing cameras typically cannot focus on the person being questioned immediately after they are called upon; the camera must wait until the person begins speaking before adjusting its view to them.
[0003] Therefore, how to improve the situation where speakers are ignored when they speak or when they designate attendees to speak is an important issue in this field. Summary of the Invention
[0004] This disclosure provides a method for composing a video frame, comprising the following steps: Obtaining a priority list, wherein the priority list contains multiple priorities for multiple person identities; Receiving multiple video streams; Identifying multiple identity markers corresponding to multiple face frames in the multiple video streams; Obtaining multiple display priorities corresponding to the multiple face frames based on the multiple identity markers and the priority list; Detecting whether the multiple face frames are speaking; Based on whether the multiple face frames are speaking and the multiple display priorities, generating a main display area of the video frame composed of at least one of the multiple face frames.
[0005] In some embodiments, the video frame composition method includes the following steps: In single-person mode, among the plurality of facial frames that are speaking, a first facial frame with a first identity marker having the highest display priority among the plurality of facial frames is determined. In this single-person mode, the first facial frame is positioned in the main display area of the video frame.
[0006] In some embodiments, the video frame composition method includes the following steps: A person with a second identity tag is designated corresponding to a person with a first identity tag. In response to the designation from the person with the first identity tag, the main display area is divided into a first split screen and a second split screen, and a question-and-answer mode is initiated. In this question-and-answer mode, a first face frame is configured on the first split screen, and a second face frame from the plurality of face frames corresponding to the second identity tag is configured on the second split screen.
[0007] In some embodiments, the video frame composition method includes the following steps: In response to the end of the question-and-answer mode, the question-and-answer mode is switched back to single-person mode to merge the first split screen and the second split screen as the main display area, and the first face frame is configured in the main display area.
[0008] In some embodiments, the designation of a person corresponding to a first identity marker is achieved by receiving a sound source signal of a person corresponding to the first identity marker via a sound receiving device, wherein the sound source signal includes a second identity marker and keywords of a question-and-answer pattern.
[0009] This disclosure provides an electronic device. The electronic device includes a storage device and processing circuitry. The processing circuitry is configured to perform the following steps: obtaining a priority list, wherein the priority list contains multiple priorities for multiple person identities; receiving multiple video streams; identifying multiple identity markers corresponding to multiple face frames in the multiple video streams; determining multiple display priorities corresponding to the multiple face frames based on the multiple identity markers and the priority list; detecting whether the multiple face frames are speaking; and generating a main display area of the video frame consisting of at least one of the multiple face frames based on whether the multiple face frames are speaking and the multiple display priorities.
[0010] In some embodiments, the processing circuitry is further configured to perform the following steps: In single-person mode, among the plurality of face frames that are speaking, determine a first face frame having the first identity marker of the highest among the plurality of display priorities. In single-person mode, position the first face frame in the main display area of the video frame.
[0011] In some embodiments, the processing circuitry is further configured to perform the following steps: Assigning a person with a second identity tag to a person corresponding to a person with a first identity tag. In response to the assignment from the person with the first identity tag, splitting the main display area into a first split screen and a second split screen and initiating a question-and-answer mode. In the question-and-answer mode, configuring a first face frame on the first split screen, and configuring a second face frame from the plurality of face frames corresponding to the second identity tag on the second split screen.
[0012] In some embodiments, the processing circuitry is further configured to perform the following steps: In response to the end of the question-and-answer mode, the question-and-answer mode is switched back to single-person mode to merge the first split screen and a second split screen as the main display area, and the first face frame is configured in the main display area.
[0013] In some embodiments, the designation of a person corresponding to a first identity marker is achieved by receiving a sound source signal of a person corresponding to the first identity marker via a sound receiving device, wherein the sound source signal includes a second identity marker and keywords of a question-and-answer pattern.
[0014] In summary, this disclosure document uses facial recognition to identify speakers with high display priority in video conferencing. Based on whether the multiple facial frames are speaking and the multiple display priorities, it determines that at least one of the multiple facial frames constitutes the main display area of the video screen, thereby enabling participants to more clearly see and notice the speaker's message. Attached Figure Description
[0015] To make the above and other objects, features, advantages and embodiments of this disclosure more apparent and understandable, the accompanying drawings are described below:
[0016] Figure 1 This is a schematic diagram of an electronic device according to an embodiment of the present disclosure;
[0017] Figure 2A This is a flowchart of a video frame composition method according to an embodiment of the present disclosure;
[0018] Figure 2B This disclosure provides an embodiment of the invention. Figure 2A The flowchart of step S170 in the process;
[0019] Figure 2C This disclosure provides an embodiment of the invention. Figure 2A The flowchart of step S140 in the middle;
[0020] Figure 3 This is a schematic diagram of an electronic device and a video stream according to an embodiment of the present disclosure;
[0021] Figure 4 This is a schematic diagram of a video stream at a point in time according to an embodiment of the present disclosure;
[0022] Figure 5 For the purpose of this disclosure, one embodiment is disclosed in relation to... Figure 4 A schematic diagram of the display screen of an electronic device at the same point in time;
[0023] Figure 6 This is a schematic diagram of a video stream at another point in time, according to an embodiment of this disclosure.
[0024] Figure 7 For the purpose of this disclosure, one embodiment is disclosed in relation to... Figure 6 A schematic diagram of the display screen of an electronic device at the same point in time;
[0025] Figure 8 This is a schematic diagram of the display screen of an electronic device according to an embodiment of the present disclosure;
[0026] Figure 9 This is a schematic diagram of the display screen of an electronic device according to an embodiment of the present disclosure.
[0027] [Symbol Explanation]
[0028] To make the above and other objects, features, advantages and embodiments of this disclosure more apparent and understandable, the accompanying symbols are explained as follows:
[0029] 100, 100a, 100b, 100c, 100d: Electronic devices
[0030] 102: Processing Circuit
[0031] 104: Storage device
[0032] 106: Radio device
[0033] 108: Photographic Device
[0034] 110: Display screen
[0035] 200: Network
[0036] 210, 220, 230, 240: Video stream
[0037] 212: First face frame
[0038] 232: Second face frame
[0039] 302: List of Participant Avatars
[0040] MA: Main display area
[0041] SC1: First Split Screen
[0042] SC2: Second Split Screen
[0043] SA1, SA2, SA3: Sub-display areas
[0044] S100: Video Screen Composition Method
[0045] Steps S110, S120, S130, S140, S142, S144, S146, S150, S160, S170, S172, S174, S176, S178, S179, S180 Detailed Implementation
[0046] The following is a detailed description of embodiments with reference to the accompanying drawings. However, the provided embodiments are not intended to limit the scope of this disclosure, and the description of the structural operation is not intended to limit the order of execution. Any structure resulting from the recombination of elements and producing a device with equivalent functionality is within the scope of this disclosure. Furthermore, the illustrations are for illustrative purposes only and are not drawn to their original dimensions. For ease of understanding, the same or similar elements will be labeled with the same symbols in the following description.
[0047] Unless otherwise specified, the terms used throughout the specification and claims generally have their ordinary meaning in the context of the art, the disclosure, and the specific content.
[0048] Furthermore, the terms "comprising," "including," "having," "containing," etc., used in this document are all open-ended terms, meaning "including but not limited to." Additionally, the term "and / or" as used in this document includes any one or more of the related listed items and all combinations thereof.
[0049] In this document, when a component is referred to as "coupled," it may mean "electrically coupled." "Coupled" can also be used to indicate the operation or interaction between two or more components. Furthermore, although terms such as "first," "second," etc., are used in this document to describe different components, these terms are only used to distinguish components or operations described using the same technical terminology.
[0050] Please see Figure 1 , Figure 1 This is a schematic diagram of an electronic device 100 according to an embodiment of this disclosure. Figure 1 As shown, the electronic device 100 includes a display screen 110, processing circuitry 102, and storage device 104. In some embodiments, the electronic device 100 may be implemented as a computer, laptop, tablet, or other device capable of receiving or transmitting video streams. The processing circuitry 102 may be implemented as a processor, microcontroller, or other component / assembly with similar functionality. The storage device 104 may be implemented as memory, cache, hard disk, or other component / assembly with similar functionality.
[0051] Electronic device 100 uses a microphone 106 to record audio or determine the direction of an audio source. Electronic device 100 uses a camera 108 to take photographs to generate a video stream. Electronic device 100 displays the video image through a display screen 110. In other embodiments, electronic device 100 may also use an external projection device for image / screen display. In some embodiments, electronic device 100 includes both a microphone and a camera. Therefore, the relative configuration of the microphone 106 and camera 108 with electronic device 100 is not limited thereto.
[0052] For a better understanding of the embodiments disclosed herein, please refer to Figure 1 , 2A ~2C and 3~7. Figure 2A This is a flowchart of a video frame composition method S100 according to an embodiment of the present disclosure. Figure 2B This disclosure provides an embodiment of the invention. Figure 2A The flowchart of step S170 in the process. Figure 2C This disclosure provides an embodiment of the invention. Figure 2AThe flowchart for step S140 is shown. The video frame composition method S100 includes steps S110 to S180. Step S170 includes steps S172 to S179. Step S140 includes steps S142 to S146. Steps S110 to S180, S142 to S149, and S172 to S179 can all be executed by the processing circuit 102 in the electronic device 100.
[0053] Figure 3 This is a schematic diagram of electronic devices 100a-100d and video streams 210, 220, 230 and 240 according to an embodiment of this disclosure. Figure 4 This is a schematic diagram of video streams 210, 220, 230, and 240 at a point in time, according to an embodiment of this disclosure. Figure 5 For the purpose of this disclosure, one embodiment is disclosed in relation to... Figure 4 A schematic diagram of the display screen 110 of the electronic device 100 at the same point in time. Figure 6 This is a schematic diagram of video streams 210, 220, 230, and 240 at another point in time, according to an embodiment of this disclosure. Figure 7 For the purpose of this disclosure, one embodiment is disclosed in relation to... Figure 6 A schematic diagram of the display screen 110 of the electronic device 100 at the same point in time.
[0054] In step S110, the meeting begins. At this time, participants from different regions / spaces who wish to attend the meeting turn on the video conferencing software on their electronic devices 100a to 100d. Figure 3 The electronic devices 100a to 100d in the middle can be made by Figure 1 The implementation of electronic device 100 is not described in detail here. Electronic devices 100a-100d respectively use audio and video recording devices to capture and record the images and audio of each meeting room, thereby generating video streams 210, 220, 230, and 240. In another embodiment, at the beginning, all participants in the meeting can be in the same meeting room / space, and electronic device 100 includes multiple video recording devices 108 and audio recording devices 106 to capture and record the meeting video / space, thereby generating multiple video streams 210, 220, 230, and 240.
[0055] In step S120, a priority list is obtained. The priority list contains multiple priorities for multiple personnel identities. For example, electronic devices 100a-100d can read / retrieve the personnel list of a company / school from the database / storage device 104. In company meetings, personnel identities can be prioritized based on job titles; for example, new employees, senior employees, department heads, managers, general managers, and chairpersons can have different priorities. In distance learning on campus, teachers can be set as higher priority and students as lower priority.
[0056] It is worth noting that the aforementioned priority of personnel identities can be achieved through facial recognition during registration. Facial features, along with the personnel's position / identity and title (name), are recorded in the database. Once the meeting begins, the identity of attendees can be determined through facial recognition without relying on a meeting account. In some embodiments, the priority of attendee identities can be set based on position. In other embodiments, the priority of attendee identities can be adjusted according to the content of the meeting.
[0057] In step S130, multiple video streams are received. The video streams 210, 220, 230, and 240 generated by electronic devices 100a to 100d are transmitted through network 200, so that each electronic device 100a to 100d can receive the video streams 210, 220, 230, and 240, thereby generating a video image and initiating the general mode of video conferencing.
[0058] In step S140, multiple identity markers corresponding to multiple face frames in the multiple video streams are identified. Please refer to [link / reference]. Figure 2C Step S140 includes steps S142 to S146.
[0059] In step S142, local images are analyzed. Furthermore, in step S144, facial recognition is used to create a meeting list. Electronic devices 100a-100d can perform facial recognition calculations on the images generated by video streams 210, 220, 230, and 240 in normal mode. In other words, electronic devices 100a-100d can compare each facial frame captured from video streams 210, 220, 230, and 240 with the facial features of individuals in the aforementioned database, thereby obtaining multiple local meeting lists. The local meeting lists include the attendees' identities (e.g., names and positions) and priorities set based on their identities.
[0060] In step S146, a list of participants from each party to the meeting is obtained. In some embodiments, electronic devices 100a-100d can transmit their respective local participant lists, calculated in step S142, to each other, thereby obtaining the identities of participants from different locations (i.e., different video spaces) in the video conference. Furthermore, processing the local video feeds using electronic devices 100a-100d in each location can save computational costs. In other instances, each of electronic devices 100a-100d can first receive all video streams 210, 220, 230, and 240, and perform facial recognition on the feeds of all video streams 210, 220, 230, and 240 to obtain the participant list. Therefore, this invention is not limited to this.
[0061] In step S150, multiple display priorities corresponding to the multiple face frames are obtained based on the multiple identity tags and the priority list. In some embodiments, the multiple display priorities of the multiple face frames can be determined by searching the priority list obtained in step S120 based on the multiple identity tags. In other embodiments, the priorities of the participants can also be set directly before the meeting starts. Therefore, this application is not limited to this.
[0062] In step S160, it is detected whether the multiple face frames are speaking. In a multi-person conference, when participants are all enthusiastically discussing the topic at the same time, the audio in the video conference may become cluttered and difficult to hear. Therefore, following step S170, based on whether the multiple face frames are speaking and the multiple display priorities, a main display area of the video screen is generated, consisting of at least one of the multiple face frames. In this way, the main display area MA of the video screen can be switched to the one who is speaking and has the highest priority among the participants, thereby reminding all participants that the supervisor, teacher, or speaker in the video conference is speaking.
[0063] Step S170 includes steps S172 to S179. Please refer to [link / reference]. Figure 2B In step S172, in single-person mode, among the multiple facial frames that are speaking, the first facial frame with the highest first identity marker among the multiple display priority frames is determined. Specifically, the electronic device 100 can use a two-dimensional array-type sound receiving device 106 to receive sounds generated from different locations in the venue and determine the direction of these sound sources, then compare the sound source directions with the video image to determine which facial frames are speaking. Next, in single-person mode, the electronic device 100 selects the facial frame with the highest display priority from the speaking facial frames.
[0064] For example, in Figure 4 In the video streams 210, 220, 230, and 240 shown, the electronic device 100 detects that the first face frame 212 is speaking, and within the speaking face frame, the first identity marker corresponding to the first face frame 212 has the highest display priority, and the first face frame 212 is positioned in the main display area MA of the video frame, such as... Figure 5 As shown.
[0065] In some embodiments, the camera device 108 has an adjustable zoom range, and the electronic device 100 can control the focal length of the camera device 108 to generate a first face frame 212 with higher resolution, and use this as the output video stream 210. In other embodiments, if the camera device 108 does not have an adjustable zoom range, the electronic device 100 can capture the first face frame 212 from the video stream 210 and enlarge the first face frame 212 as the output video stream 210.
[0066] exist Figure 5 In the illustrated embodiment, in single-person mode, the video screen includes a main display area MA and sub-display areas SA1 to SA3. Sub-display areas SA1 to SA3 are used to configure video streams 220, 230, and 240, respectively. In other embodiments, the video screen in single-person mode may not have sub-display areas SA1 to SA3; the video screen may consist only of the main display area MA, thereby allowing for a clearer view of the speaker, supervisor, or teacher.
[0067] In step S174, the person corresponding to the first identity tag designates the person with the second identity tag. In embodiments of this disclosure, the person with the first identity tag can designate the person with the second identity tag.
[0068] In some embodiments disclosed herein, the designation of a person corresponding to the first identity marker can be achieved by the receiving device 106 receiving the sound source signal of the person corresponding to the first identity marker. The sound source signal includes the second identity marker (e.g., the title or name of the designated person) and the starting keyword of the question-and-answer mode. For example, the sound source signal received by the receiving device 106 from the sound source direction of the first face frame 212 corresponding to the first identity marker within a predetermined time, after being identified by the electronic device 100, yields the words "answer" and "Elsa." Therefore, the person with the first identity marker can designate the person with the second identity marker by voice control or by selecting the avatar of the designated person (e.g., the list of attendee avatars 302 presented on the display screen 110), or by directly selecting the face of the attendee in the video frame. Step S176 then proceeds.
[0069] In step S176, the main display area (such as...) Figure 5 The main display area (MA) shown is broken down into the first segmented screen (e.g., Figure 5 The first split screen SC1 and the second split screen SC2 shown are as follows: Figure 5 The second split screen (SC2) is shown, and the question-and-answer mode begins, as shown. Figure 7 As shown.
[0070] In step S178, in question-and-answer mode, the first face frame 212 is configured on the first split screen SC1, and the second face frame 232 is configured on the second split screen SC2. This shortens the time spent searching for the speaker or designated person during the question-and-answer process in a video conference, and timely switching of the video screen can also remind the designated person to answer the question. In some embodiments, when the processing circuit 102 receives an instruction to designate a second-identified person, the edge area of the display screen 110 of the electronic device 100 where the designated second-identified person is located can flash as a prompt, thereby facilitating the meeting process.
[0071] For example, in Figure 6 In the video streams 210, 220, 230, and 240 shown, electronic device 100 detects that a person corresponding to a first identity marker in a first facial frame 212 is designating a person with a second identity marker (e.g., Elsa). Figure 7 In the video frame shown, a first face frame 212 is configured in the first split frame SC1, and a second face frame 232 corresponding to the second identity marker (e.g., Elsa) is configured in the second split frame SC2.
[0072] Similarly, the second face frame 232 can be sensed by the electronic device 100 controlling the focal length of the camera device 108, or captured, magnified and output by the electronic device 100 from the video stream 230.
[0073] In step S179, in response to the end of the question-and-answer mode, the video feed is switched back from question-and-answer mode to single-person mode. Specifically, in response to the end of the question-and-answer mode, the switch back to single-person mode merges the first split screen SC1 and the second split screen SC2 into the main display area MA, and the first face frame 212 corresponding to the first identity marker is configured in the main display area MA. Following step S174, another participant is designated to continue the question-and-answer mode, or step S180 is performed, ending the meeting.
[0074] Figure 8 This is a schematic diagram of the display screen 110 of an electronic device 100 according to an embodiment of the present disclosure. Figure 9 This is a schematic diagram of the display screen 110 of an electronic device 100 according to an embodiment of this disclosure. Compared to Figure 5 as well as Figure 7 Multiple attendees in the same venue used the same set of audio equipment. Figure 8 as well as Figure 9 In this scenario, when the video stream output shows only one person, it can directly detect video streams with sound, determine which person has the highest display priority, and switch to that person accordingly. Figure 8 The single-player mode shown or Figure 9 The question-and-answer format shown. Figure 8 as well as Figure 9 The remaining operations of the embodiments described above are similar to those described above. Figure 5 as well as Figure 7 Therefore, the operation can also be performed by steps S110 to S180, as in the embodiment.
[0075] In summary, the disclosed embodiments, through pre-registered personnel identities and corresponding facial features, identify speakers with high display priority in single-person video conferencing via facial recognition. When the speaker speaks, their facial frame is positioned in the main display area MA of the video screen, allowing participants to clearly see and notice the speaker's important information. Furthermore, when a speaker with high display priority designates a participant for a question-and-answer session, the electronic device 100 can directly respond to the designation by switching the video screen to question-and-answer mode, simultaneously displaying the facial frames of both the speaker (questioner) and the designated person. This makes the video conferencing process smoother. By switching the meeting screen, participants can avoid having to identify the speaker or designated person based on their voice, allowing each participant to synchronize with the meeting progress.
[0076] Although the present disclosure has been described above with reference to embodiments, it is not intended to limit the present disclosure. Any person skilled in the art may make various modifications and refinements without departing from the spirit and scope of the present disclosure. Therefore, the scope of protection of the present disclosure shall be determined by the scope defined in the appended claims.
Claims
1. A method for composing video frames, characterized in that, Include: Obtain a priority list, which contains multiple priorities for multiple personnel identities; Receive multiple video streams from multiple meeting spaces, wherein at least one of the multiple video streams contains multiple face frames of multiple participants in the same meeting space; Identify multiple identity markers corresponding to the multiple facial frames in the multiple video streams; Based on the multiple identity markers and the priority list, multiple display priorities corresponding to the multiple face frames are obtained; Detecting whether the multiple facial frames are speaking includes: using a two-dimensional array of sound receiving devices to receive sound in the same conference space and determine the directions of multiple sound sources, comparing the directions of the multiple sound sources with the video frame of at least one of the multiple video streams, thereby determining at least one of the multiple facial frames is speaking; Based on whether the plurality of facial frames speak and the plurality of display priorities, a main display area for a video frame is generated, which is composed of at least one of the plurality of facial frames; Configure the images of the multiple video streams in multiple sub-display areas of the video frame; In a single-person mode, among the multiple face frames that are speaking, a first face frame with the highest first identity marker among the multiple display priority frames is determined; as well as In the single-player mode, the first face frame is positioned in the main display area of the video image.
2. The video frame composition method according to claim 1, characterized in that, Include: In response to the designation from the person with the first identity tag, a person with a second identity tag is designated, and the main display area is split into a first split screen and a second split screen and a question-and-answer mode is started. as well as In this question-and-answer mode, the first face frame is configured on the first segmented screen, and a second face frame corresponding to the second identity marker is configured on the second segmented screen.
3. The video frame composition method according to claim 2, characterized in that, Include: In response to the end of the question-and-answer mode, the question-and-answer mode is switched back to the single-player mode to merge the first split screen and a second split screen into the main display area, and the first face image frame is configured in the main display area.
4. The video frame composition method according to claim 2, characterized in that, The designation of the person corresponding to the first identity mark is achieved by receiving a sound source signal from the person corresponding to the first identity mark via a radio device, wherein the sound source signal contains the second identity mark and keywords of the question-and-answer mode.
5. An electronic device, characterized in that, Include: A storage device; and A processing circuit, used to: Obtain a priority list, which contains multiple priorities for multiple personnel identities; Receive multiple video streams, wherein at least one of the multiple video streams contains multiple face frames of multiple participants in the same meeting space; Identify multiple identity markers corresponding to the multiple facial frames in the multiple video streams; Based on the multiple identity markers and the priority list, multiple display priorities corresponding to the multiple face frames are obtained; Detecting whether the multiple facial frames are speaking includes: using a two-dimensional array of sound receiving devices to receive sound in the same conference space and determine the directions of multiple sound sources, comparing the directions of the multiple sound sources with the video frame of at least one of the multiple video streams, thereby determining at least one of the multiple facial frames is speaking; Based on whether the plurality of facial frames speak and the plurality of display priorities, a main display area for a video frame is generated, which is composed of at least one of the plurality of facial frames; In the plurality of face frames that are currently speaking, determine a first face frame that has the highest first identity marker among the plurality of display priorities; The main display area is divided into a first segmented screen and a second segmented screen, and a question-and-answer mode is started. as well as In the question-and-answer mode, the first face frame is configured on the first segmented screen, and a second face frame corresponding to a second identity tag is configured on the second segmented screen.
6. The electronic device according to claim 5, characterized in that, This processing circuit is further used for: In a single-player mode, the first face frame is positioned in the main display area of the video frame.
7. The electronic device according to claim 6, characterized in that, This processing circuit is further used for: The person corresponding to the first identity tag designates the person with the second identity tag.
8. The electronic device according to claim 7, characterized in that, This processing circuit is further used for: In response to the end of the question-and-answer mode, the question-and-answer mode is switched back to the single-player mode to merge the first split screen and a second split screen into the main display area, and the first face image frame is configured in the main display area.
9. The electronic device according to claim 7, characterized in that, The designation of the person corresponding to the first identity mark is achieved by receiving a sound source signal from the person corresponding to the first identity mark via a radio device, wherein the sound source signal contains the second identity mark and keywords of the question-and-answer mode.
Citation Information
Patent Citations
Video conference control method and system, mobile terminal and storage medium
CN113139491A
Information processing apparatus, conference system, information processing method, and program
JP2017092675A
Video conference system, especially for automatically selecting an opposite party through voice recognition according to a priority order
KR1019990060724A