A multi-user screen mirroring method and system

By obtaining the correspondence between the projection signal and the user's identity, and by using voice recognition and keyboard and mouse actions to automatically switch the projection signal, the problem of low efficiency in switching projection content in existing technologies is solved, and the intelligence and efficiency of multi-user projection are improved.

CN115421679BActive Publication Date: 2026-04-03EXANDS INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-02
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, users need to manually switch between projected content, which is inefficient and cannot meet the needs of multi-user screen sharing scenarios.

Method used

By acquiring the correspondence between multiple screen projection signals and user identities, and utilizing user voice information and keyboard and mouse action recognition, the screen projection signal is automatically switched to achieve intelligent switching of screen projection content.

Benefits of technology

It enables intelligent switching of screen-casting content, improving the efficiency and user experience of multi-user screen casting and reducing the time spent on manual operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115421679B_ABST
    Figure CN115421679B_ABST
Patent Text Reader

Abstract

This specification provides a multi-user screen casting method and system. The method includes acquiring multiple screen casting signals, each of which corresponds to the user identity of a user among the multiple users; acquiring user voice information of at least one user among the multiple users; and determining a casting method for the multiple screen casting signals based on the user voice information, the casting method including switching between the multiple screen casting signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This manual relates to the field of smart devices, and in particular to a multi-user screen mirroring method and system. Background Technology

[0002] Projection devices are widely used large-screen imaging devices that project content from terminal devices onto a large screen for wider viewing and to enhance the user experience. Common projection devices connect to a single terminal device, projecting its content. With the accelerating pace of modern work and the increasing prevalence of shared screens (such as multi-person meetings), new demands are being placed on how projection content is switched. For example, the ability to intelligently switch projection content based on user needs is required. Currently, users must manually switch projection content, a time-consuming and inefficient process.

[0003] Therefore, we hope to propose a multi-user screen casting method and system for intelligently switching screen casting content, thereby effectively solving the problem of manually switching screen casting content. Summary of the Invention

[0004] This specification provides one or more embodiments of a multi-user screen casting method. The multi-user screen casting method includes acquiring multiple screen casting signals, each of which corresponds to a user identity among the multiple users; acquiring user voice information of at least one of the multiple users; and determining a casting method for the multiple screen casting signals based on the user voice information, the casting method including switching between the multiple screen casting signals.

[0005] This specification provides one or more embodiments of a multi-user screen projection system. The system includes a first acquisition module for acquiring multiple projection signals, each projection signal corresponding to a user identity among multiple users; a second acquisition module for acquiring user voice information of at least one of the multiple users; and a first determination module for determining a projection method for the multiple projection signals based on the user voice information, the projection method including switching between the multiple projection signals.

[0006] In some embodiments, the system further includes a third acquisition module for acquiring the keyboard and mouse actions of the current speaker; a judgment module for judging whether the keyboard and mouse actions are applied to the screen projection application; and a second determination module for determining the projection method of multiple screen projection signals in response to the keyboard and mouse actions being applied to the screen projection application.

[0007] This specification provides one or more embodiments of a multi-user screen mirroring device. The device includes at least one storage medium storing computer instructions; and at least one processor executing the computer instructions to implement the multi-user screen mirroring method.

[0008] This specification provides one or more embodiments of a computer-readable storage medium that stores computer instructions. When a computer reads the computer instructions from the storage medium, the computer executes a multi-user screen mirroring method. Attached Figure Description

[0009] This specification will be further described by way of exemplary embodiments, which will be described in detail with reference to the accompanying drawings. These embodiments are not limiting; in these embodiments, the same reference numerals denote the same structures, wherein:

[0010] Figure 1 These are schematic diagrams illustrating application scenarios of a multi-user screen mirroring system according to some embodiments of this specification;

[0011] Figure 2 This is a block diagram of a multi-user screen mirroring system according to some embodiments of this specification;

[0012] Figure 3 This is an exemplary flowchart of a multi-user screen casting method according to some embodiments of this specification;

[0013] Figure 4 This is a schematic diagram of the feature extraction model shown in some embodiments of this specification;

[0014] Figure 5 This is a schematic diagram of the structure of a text judgment model according to some embodiments of this specification;

[0015] Figure 6 This is an exemplary flowchart of a multi-user screen mirroring method according to some embodiments of this specification. Detailed Implementation

[0016] To more clearly illustrate the technical solutions of the embodiments in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely some examples or embodiments of this specification. For those skilled in the art, these drawings can be applied to other similar scenarios without creative effort. Unless obvious from the context or otherwise specified, the same reference numerals in the drawings represent the same structures or operations.

[0017] It should be understood that the terms “system,” “device,” “unit,” and / or “module” used herein are one way to distinguish different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.

[0018] As indicated in this specification and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of expressly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0019] Flowcharts are used in this specification to illustrate the operations performed by the system according to embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the steps can be processed in reverse order or simultaneously. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0020] Figure 1 This is a schematic diagram illustrating the application scenario of a multi-user screen mirroring system according to some embodiments of this specification.

[0021] like Figure 1 As shown, the screen mirroring switching system 100 may include user 110, user computer information 120, server 130, screen mirroring content before switching 140, and screen mirroring content after switching 150.

[0022] User 110 is a user who may need to cast their screen. For example... Figure 1 As shown, user 110 may include multiple users such as user 110-1, user 110-2, and user 110-3.

[0023] User computer information 120 refers to information stored on the user's computer, such as spreadsheets, documents, and PowerPoint presentations. Figure 1 As shown, users 110-1, 110-2, and 110-3 correspond to user computer information 120-1, user computer information 120-2, and user computer information 120-3, respectively.

[0024] Server 130 can be used to manage resources and process data and / or information from at least one component of the screen mirroring switching system 100 or an external data source (e.g., a cloud data center). In some embodiments, server 130 can determine the screen mirroring content 140 before switching based on the user's voice information. See the relevant description in the flowchart for details on determining the screen mirroring content 140 before switching.

[0025] The content projected before switching (140) refers to the content projected by the projection device before switching.

[0026] The switched projection content 150 refers to the projection content after the projection device switches to the previous projection content. In some embodiments, the server 130 can determine whether to switch the previous projection content 140 to the switched projection content 150 based on the user's voice information.

[0027] The components of the screen mirroring switching system 100 can achieve communication connections through wired connections, wireless connections, or a combination thereof.

[0028] Figure 2 This is a block diagram of a multi-user screen mirroring system according to some embodiments of this specification. For example... Figure 2 As shown, system 200 includes the following modules.

[0029] In some embodiments, the first acquisition module 210 can be used to acquire multiple screen projection signals, and each of the multiple screen projection signals corresponds to the user identity of each of the multiple users. For details on how to acquire multiple screen projection signals, please refer to [link to relevant documentation]. Figure 3 And related information.

[0030] In some embodiments, the second acquisition module 220 may be used to acquire user voice information of at least one of a plurality of users.

[0031] In some embodiments, the first determining module 230 can be used to determine the projection method for multiple projection signals based on user voice information, wherein the projection method includes switching between multiple projection signals. For details on how to determine the projection method for multiple projection signals based on user voice information, please refer to... Figure 3 And related information.

[0032] In some embodiments, the first determining module 230 may be used to determine at least one user identity corresponding to the user's voice information through voice recognition based on the user's voice information; determine a target projection signal based on the correspondence between at least one user identity and multiple projection signals; and determine a switching instruction for the projection signal of the projection device based on the target projection signal.

[0033] In some embodiments, the first determining module 230 can be used to determine the user identity of the current speaker based on the user voice information of the current speaker; convert the user voice information into text information; determine the probability that the text information is a declarative sentence based on the text information through a text judgment model; and control the projection device to switch the projection signal corresponding to the user identity when the probability meets the preset conditions.

[0034] In some embodiments, the system 200 may further include a third acquisition module 240, a judgment module 250, and a second determination module 260.

[0035] In some embodiments, the third acquisition module 240 can be used to acquire the keyboard and mouse actions of the current speaker.

[0036] In some embodiments, the determination module 250 can be used to determine whether keyboard and mouse actions are applied to the screen mirroring application.

[0037] In some embodiments, the second determining module 260 can be used to determine the projection method of multiple projection signals in response to keyboard and mouse actions acting on the projection application. For details on how to determine the projection method of multiple projection signals in response to keyboard and mouse actions acting on the projection application, please refer to... Figure 6 And related information.

[0038] It should be noted that the above description of the screen mirroring switching system and its modules is for convenience only and should not be construed as limiting this specification to the scope of the illustrated embodiments. It is understood that those skilled in the art, after understanding the principles of this system, may arbitrarily combine the various modules or construct subsystems connected to other modules without departing from these principles. In some embodiments, Figure 2 The first acquisition module, second acquisition module, first determination module, third acquisition module, judgment module, and second determination module disclosed herein can be different modules within a single system, or a single module can implement the functions of two or more of the aforementioned modules. For example, the modules can share a single storage module, or each module can have its own separate storage module. Such variations are all within the scope of protection of this specification.

[0039] Figure 3 This is an exemplary flowchart illustrating a multi-user screen mirroring method according to some embodiments of this specification. Figure 3 As shown, process 300 includes the following steps. In some embodiments, process 300 may be executed by server 130.

[0040] Step 310: Obtain multiple screen projection signals, and each of the multiple screen projection signals corresponds to the user identity of each of the multiple users.

[0041] A screen mirroring signal is a signal associated with the content a user mirrors. In some embodiments, the content a user needs to mirror has a corresponding screen mirroring signal, meaning there is a correspondence between the screen mirroring signal and the user's identity. For example, if user 1, user 2, and user 3 need to mirror content 1, content 2, and content 3 respectively, where content 1, content 2, and content 3 correspond to screen mirroring signals 1, 2, and 3 respectively, then screen mirroring signals 1, 2, and 3 correspond to user 1, user 2, and user 3 respectively. The correspondence between the screen mirroring signal and the user's identity can be based on a preset value, such as by registering and reporting to determine the correspondence. The correspondence between the screen mirroring signal and the user's identity can be pre-stored in a storage device, and this correspondence can be obtained by the first acquisition module.

[0042] In some implementations, the first acquisition module can be used to acquire the projection signal. The first acquisition module can acquire the projection signal via wired, wireless, or a combination thereof. In some embodiments, the first acquisition module can also invoke the projection signal pre-stored in a storage device.

[0043] Step 320: Obtain user voice information of at least one of the multiple users.

[0044] User voice information can include the user's timbre, tone of voice, and content of speech. Timbre can be used to determine the correspondence between the user and the projection signal; for example, if the speaker is user 1, then the corresponding projection signal is projection signal 1. Tone of voice and content of speech can both be used to determine whether the user needs to switch the projection content.

[0045] In some implementations, user voice information can be acquired through a second acquisition module. In some implementations, user voice information can be acquired through an input device of the terminal device, such as a microphone.

[0046] Step 330: Determine the projection method for multiple projection signals based on the user's voice information. The projection method includes switching between multiple projection signals.

[0047] The casting method refers to the switching between multiple casting signals, or the maintenance or adjustment of the casting method under the same casting signal. The casting method may include whether to switch casting signals, adjust the current casting application window, or any other adjustment to the casting state; for example, zooming in or out of the casting application, switching casting applications, etc.; or, for example, switching from casting signal 1 to casting signal 2 while maintaining the casting content of casting signal 1. In some embodiments, the first determining module can be used to determine the casting method for multiple casting signals based on user voice information. For example, the first determining module can determine whether to switch the current casting signal based on user voice information. If the user voice information explicitly contains content about switching casting signals, then the current casting signal is switched; see [link to documentation] for details. Figure 6 The corresponding description. In some embodiments, the first determining module can determine the projection method for multiple projection signals through a feature extraction model and a text judgment model; for details, please refer to [link to relevant documentation]. Figure 4 and Figure 5 The relevant description is provided. In some embodiments, the first determining module can also determine the projection method for multiple screen projection signals by judging keyboard and mouse actions; for details, please refer to [link to relevant documentation]. Figure 6 Related descriptions.

[0048] It should be noted that the above description of process 300 is for illustrative purposes only and does not limit the scope of this specification. Those skilled in the art can make various modifications and changes to process 300 under the guidance of this specification. However, these modifications and changes remain within the scope of this specification.

[0049] In some embodiments, at least one user identity corresponding to user voice information can be determined through speech recognition based on user voice information. Speech recognition may include recognizing user voice information using a feature extraction model and determining at least one user identity based on the recognition result.

[0050] Figure 4 This is a schematic diagram of the feature extraction model shown in some embodiments of this specification.

[0051] like Figure 4 As shown, the input of the feature extraction model 460 includes user voice information C 450, and the output includes user identity 470.

[0052] In some embodiments, the feature extraction model 460 may be a deep neural network model, a recurrent neural network model, etc.

[0053] like Figure 4 As shown, the feature extraction model 460 includes a speech feature extraction layer C 460-1 and a feature matching layer 460-2.

[0054] In some embodiments, the speech feature extraction layer can determine speech feature vectors based on user speech information.

[0055] In some embodiments, the feature matching layer 460-2 can determine the user's identity 470 based on the speech feature vector C 430-3 and the reference feature vector. For example, if the vector similarity between the speech feature vector C 430-3 and the reference feature vector is greater than a preset value, then the user's identity is confirmed as the user corresponding to the reference feature vector. The vector similarity can be determined by calculating the distance between the two vectors.

[0056] The reference feature vector is a pre-extracted speech feature vector of a known user, used to compare with the speech feature vector extracted by the speech feature extraction layer to determine the user's identity. For example, if the similarity between the reference feature vector of user 1 and the speech feature vector extracted by the speech feature extraction layer is greater than a preset value, then the user identity corresponding to the user's speech information is user 1.

[0057] In some embodiments, the feature extraction model can be trained from a consistency model. For more information on the training process of the feature extraction model, please refer to [link to relevant documentation]. Figure 5 Content related to the consistency model.

[0058] In some embodiments, the feature extraction model can be trained using a consistency model. The consistency model 420 includes a speech feature extraction layer A 420-1, a speech feature extraction layer B 420-2, and a judgment layer 420-3. The input to the consistency model is user speech information A 410-1 and user speech information B 410-2, and the output of the consistency model is whether the two user speech information belong to the same user 440.

[0059] In some embodiments, the speech feature extraction layer can determine speech feature vectors A 430-1 and B 430-2 based on user speech information A 410-1 and user speech information B 410-2, respectively. The judgment layer 420-3 determines whether user speech information A 410-1 and user speech information B 410-2 belong to the same person by comparing speech feature vectors A 430-1 and B 430-2.

[0060] In some embodiments, speech feature extraction layer A 420-1 and speech feature extraction layer B 420-2 can be convolutional neural network models. In some embodiments, decision layer 420-3 can be a neural network model.

[0061] A consistency model can be trained using multiple sets of labeled training samples. These training samples can include multiple sets of voice information from two users. The two user voice information sets may or may not belong to the same person. When the two user voice information sets belong to the same person, the label is "yes"; when they do not belong to the same person, the label is "no". Multiple sets of labeled training samples can be input into the consistency model. A loss function is constructed based on the output of the consistency model and the labels. The parameters of the consistency model are iteratively updated based on the loss function until the loss function meets a preset condition. For example, the loss function converges, or the loss function value is less than a preset value. When the loss function meets the preset condition, the model training is complete, and the trained consistency model is obtained.

[0062] In some embodiments, the feature extraction model can be transferred to a pre-trained consistency model. For example, the speech feature extraction layer in the pre-trained consistency model can be used as the speech feature extraction layer in the feature extraction model to determine the user's identity.

[0063] If a feature extraction model is trained alone, the training sample labels are speech feature vectors, which are not easy to obtain, i.e., not easy to manually label. Therefore, the feature extraction model can be based on establishing a consistency model for transfer learning, thereby improving the training speed of the feature extraction model, reducing the cost of manual labeling, and ensuring the accuracy of the trained feature extraction model.

[0064] In some embodiments, a target projection signal can be determined based on the correspondence between at least one user identity and multiple projection signals.

[0065] The target projection signal is the projection signal corresponding to a known user. For example, if a user's voice information is identified as user 1 through voice recognition, then the projection signal corresponding to user 1 is the target projection signal.

[0066] In some embodiments, the first determining module can match user identities with projection signals through an claiming process to determine the target projection signal. The claiming process refers to the process of associating projection signals with user identities. For example, the claiming process includes associating user 1, user 2, and user 3 with projection signal 1, projection signal 2, and projection signal 3, respectively, to determine the target projection signal.

[0067] In some embodiments, users can claim a corresponding screen projection signal based on voice recognition technology. For example, a user can claim a screen projection signal by voice inputting "This is my screen projection signal." In some embodiments, users can claim a corresponding screen projection signal based on preset data on the terminal device. For example, a user can claim a corresponding screen projection signal based on user voice information and other data already recorded on the terminal device. That is, if the terminal device stores user 1's voice information, then the screen projection signal on the terminal device corresponds to user 1.

[0068] In some embodiments, the switching instruction for the projection signal of the projection device can be determined based on the target projection signal.

[0069] A switching command is an instruction to switch the content to be cast, such as an instruction to switch the casting signal from one content to another content to be cast.

[0070] In some embodiments, the switching command can be obtained via a remote control, an app, or other means. For example, a remote control can be used to control the screen projection signal to be switched, or the screen projection signal to be switched can be selected in the app. In some embodiments, the switching command can be determined based on the user identity identified by a feature extraction model. For example, if the feature extraction model identifies the user's voice information as user 1, and user 1 corresponds to the target screen projection signal, then the switching command can be determined to be switching the screen projection signal to the target screen projection signal.

[0071] In some embodiments, the user identity of the current speaker can be determined based on the user's voice information. For example, the feature extraction model described above can be used to determine the identity of the current speaker.

[0072] In some embodiments, user voice information can be converted into text information, for example, by using speech recognition technology.

[0073] In some embodiments, the likelihood of the text information being a declarative sentence can be determined based on the text information using a text judgment model.

[0074] Figure 5 This is a structural schematic diagram of a text judgment model according to some embodiments of this specification.

[0075] like Figure 5 As shown, the input of the text judgment model 520 is text information 510, and the output of the text judgment model 520 is the probability 530 that the text is a declarative sentence.

[0076] Text information can be text generated by converting user speech information through speech recognition technology. The probability that the text is a declarative sentence can be represented by a probability threshold, for example, the threshold can be set to a number between 0 and 1, where 0 means that the probability of the text being a declarative sentence is 0, and 1 means that the text is a declarative sentence.

[0077] In some embodiments, the text judgment model can be a deep neural network model, a recurrent neural network model, etc.

[0078] The initial text judgment model can be trained based on training samples and their labels. The initial text judgment model can be a text judgment model 520 without set parameters. Training samples can include text information, and labels can be the probability that the text is a declarative sentence. Training samples can be determined based on historical user voice information, and labels can be manually labeled. The training samples are input into the initial text judgment model, and a loss function is constructed based on the output of the initial text judgment model and the labels. The parameters of the initial text judgment model are iteratively updated based on the loss function until a preset condition is met, completing the training and obtaining the trained text judgment model 520. The preset condition can be that the loss function is less than a threshold, convergence, or the training period reaches a threshold.

[0079] When the probability of the text being a declarative sentence meets the condition, the module determines to perform a screen switching operation. For example, if the current projection signal corresponds to user 1, and after user 1 finishes speaking, user 2 says, "Now I'll say something," the text judgment model determines that the probability of the text being a declarative sentence is 1, so the module determines to switch the screen to user 2's projection signal; if user 2 says, "I have a question, what is the main point of the first page you wrote?", the text judgment model determines that the probability of the text being a declarative sentence is 0, so no screen switching occurs in this case.

[0080] In some embodiments, natural language processing (NLP) techniques can be used for comprehensive judgment. If the text judgment model determines that the probability of the text being a declarative sentence is 0, but the text contains specific content, a screen switching prompt can be initiated for the user to decide whether to switch screens. The specific content can be specific words; for example, in the text "Can the next person speak?", the word "send" is a specific word. In this case, although the probability of the text being a declarative sentence is 0, the user needs to decide whether to switch screens. Specific words can also include words such as "finished speaking" or "end". Combining NLP techniques for comprehensive judgment can effectively reduce the probability of misjudgment and increase the accuracy of text judgment.

[0081] In some embodiments, the first determining module sends a prompt to switch screens before performing the screen-switching operation. The prompt is used for user confirmation of whether to switch screens. For example, the prompt may be a pop-up window displayed on the terminal device, where the user selects whether to switch screens. In some embodiments, the user can recognize the user's confirmation action through a camera. The user confirmation action includes a nodding confirmation action or a gesture confirmation action, where the gesture confirmation action may be a thumbs-up, an OK gesture, etc.

[0082] In some embodiments, the training data for the text judgment model can be user voice information judged by the feature extraction model. This user voice information is judged by the feature extraction model to be a switching instruction for the screen projection signal of the projection device. Because the probability that the text information corresponding to the user voice information is a declarative sentence meets a preset condition, the user voice information, after being converted into text information, can be used for model training of the text judgment model. For example, user voice information A can determine the switching instruction for the screen projection signal of the projection device after being recognized by the feature extraction model. After user voice information A is converted into text information a, the probability that the text information a is a declarative sentence is 1, and its corresponding label is relatively clear (the label is that the probability that text information a is a declarative sentence is 1). Then, text information a can be used as a training sample for the text judgment model.

[0083] Using user speech information that has been judged in the feature extraction model to train the text judgment model can effectively reduce the model training cost and improve training efficiency.

[0084] like Figure 5 As shown, the probability 530 that the text is a declarative sentence is related to keyboard and mouse actions 540. For details on keyboard and mouse actions, see [link to documentation]. Figure 6 And its related descriptions.

[0085] In some embodiments, if keyboard and mouse actions are applied to the screen mirroring application, the probability that the text is a declarative sentence is increased. For example, if the probability of the text being a declarative sentence based on the text information is 0.7, then when keyboard and mouse actions are applied to the screen mirroring application, the probability of the text being a declarative sentence is increased by 0.2, meaning the probability of the text being a declarative sentence is 0.7 + 0.2 = 0.9. For details on screen mirroring applications, please refer to [link to relevant documentation]. Figure 6 And its related descriptions.

[0086] Introducing keyboard and mouse actions during the training process of a text judgment model can improve the accuracy of its predictions. This makes the text judgment model more realistic in application, and sending a prompt before switching screens can effectively prevent accidental screen switching.

[0087] In some embodiments, in response to a preset condition being met by the probability that the text is a declarative sentence, the screen projection device is controlled to switch the screen projection signal corresponding to the user's identity. For example, when the probability that the text is a declarative sentence is greater than 0.8, the screen projection device is controlled to switch the screen projection signal corresponding to the user's identity. The switching operation may include switching the screen projection content.

[0088] Figure 6 This is an exemplary flowchart of a multi-user screen mirroring method according to some embodiments of this specification.

[0089] like Figure 6 As shown, process 600 includes the following steps. In some embodiments, process 600 may be executed by server 130.

[0090] Step 610: Obtain the keyboard and mouse actions of the current speaker.

[0091] Keyboard and mouse actions can include clicking the mouse, pressing at least one key, and other operations. For example, clicking the mouse to open PowerPoint, pressing a key to switch Word to the current window, clicking to exit the current window, or pressing the ESC key to exit.

[0092] For details on verifying the current speaker's user identity, please refer to [link / reference]. Figure 4 And related descriptions. In some embodiments, the current speaker's keyboard and mouse actions can be acquired by the first acquisition module in association with the screen projection signal. In some embodiments, the current speaker's keyboard and mouse actions can be acquired through the input device of the terminal device, such as a mouse, keyboard, or other input device. In some embodiments, the third acquisition module can acquire the current speaker's keyboard and mouse actions. The third acquisition module transmits the acquired keyboard and mouse actions of the current speaker to the judgment module for processing.

[0093] Step 620: Determine whether the keyboard and mouse actions are applied to the screen mirroring application.

[0094] The judgment module can be used to determine whether keyboard and mouse actions are applied to the casting application. A casting application refers to an application that needs to be projected for display; for example, casting applications may include PowerPoint, Word, and other applications. In some embodiments, the casting application can be pre-specified; for example, the user can pre-specify PowerPoint, Word, or other applications as the casting application. In some embodiments, the casting application can be the foreground window from the last time it was projected.

[0095] Step 630: In response to keyboard and mouse actions acting on the screen mirroring application, determine the projection method of multiple screen mirroring signals.

[0096] If keyboard and mouse actions are applied to the screen mirroring application, the second determining module can be used to determine the projection method of multiple screen mirroring signals. In some embodiments, if the current speaker has user voice information and also has keyboard and mouse actions on the screen mirroring application, a screen mirroring signal switching operation is performed, which includes switching the screen mirroring content. In this case, the current speaker's user voice information has been determined as a declarative sentence by the text judgment model. For example, if user 1's voice information is "Look at my PPT," and the user clicks the mouse to open the PPT, then the screen mirroring content is switched to user 1's. In some embodiments, if the application displayed in the currently foreground window is not the screen mirroring application specified by the user, then no screen switching occurs.

[0097] In some embodiments, the screen-switching operation may further include a prompt to switch screens. The prompt to switch screens is used to ask the user to confirm whether to switch screens. The prompt to switch screens may be in the form of a pop-up window or similar method.

[0098] In some embodiments, the camera can be used to identify whether the current speaker is looking up at the projected screen before sending a prompt to switch screens, thus determining whether to send a prompt. In some embodiments, if the total time the current speaker spends looking up at the projected screen during their speech exceeds a preset threshold, the system will directly switch to the projected screen signal corresponding to the preset threshold.

[0099] By identifying whether the user speaking is looking at the projected screen, the system can determine whether to send a screen-switching prompt. This improves the accuracy of screen-switching judgment. When the time it is determined that the user speaking is looking at the projected screen for more than a preset threshold, the screen can be switched directly, which can improve the efficiency of screen-switching and thus ensure user experience.

[0100] In some embodiments, a multi-user screen mirroring device may include at least one storage medium and at least one processor. The storage medium stores computer instructions, and the processor executes the computer instructions to implement the multi-user screen mirroring method.

[0101] In some embodiments, the storage medium may store computer instructions, and when the computer reads the computer instructions, the computer executes a multi-user screen mirroring method.

[0102] The basic concepts have been described above. Obviously, for those skilled in the art, the detailed disclosure above is merely illustrative and does not constitute a limitation of this specification. Although not explicitly stated herein, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are suggested in this specification and therefore remain within the spirit and scope of the exemplary embodiments described herein.

[0103] Furthermore, this specification uses specific terms to describe embodiments thereof. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic associated with at least one embodiment of this specification. Therefore, it should be emphasized and noted that references to "an embodiment," "one embodiment," or "an alternative embodiment" in different locations throughout this specification do not necessarily refer to the same embodiment. Moreover, certain features, structures, or characteristics in one or more embodiments of this specification can be appropriately combined.

[0104] Furthermore, unless expressly stated in the claims, the order of processing elements and sequences, the use of numbers and letters, or other names described in this specification are not intended to limit the order of the processes and methods described herein. Although various examples have been discussed in the foregoing disclosure of some embodiments of the invention that are currently considered useful, it should be understood that such details are for illustrative purposes only, and the appended claims are not limited to the disclosed embodiments; rather, the claims are intended to cover all modifications and equivalent combinations that conform to the spirit and scope of the embodiments described herein. For example, while the system components described above can be implemented using hardware devices, they can also be implemented solely using software solutions, such as installing the described system on existing servers or mobile devices.

[0105] Similarly, it should be noted that, in order to simplify the description disclosed herein and thus aid in the understanding of one or more embodiments of the invention, the foregoing description of embodiments in this specification may sometimes combine multiple features into a single embodiment, drawing, or description thereof. However, this method of disclosure does not imply that the subject matter of this specification requires more features than those mentioned in the claims. In fact, the embodiments contain fewer features than all the features of a single embodiment disclosed above.

[0106] In some embodiments, numbers describing the quantity of components and attributes are used. It should be understood that such numbers used in the description of embodiments are modified in some examples with the terms "approximately," "approximately," or "generally." Unless otherwise stated, "approximately," "approximately," or "generally" indicates that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may be changed depending on the characteristics required by individual embodiments. In some embodiments, numerical parameters should take into account specified significant digits and employ a general method of digit reservation. Although the numerical ranges and parameters used to confirm their breadth of range in some embodiments of this specification are approximate values, in specific embodiments, such values ​​are set as precisely as feasible.

[0107] For each patent, patent application, patent application publication, and other material, such as articles, books, specifications, publications, and documents, referenced in this specification, the entire contents of which are incorporated herein by reference. This excludes historical application documents that are inconsistent with or conflict with the content of this specification, as well as documents that limit the broadest scope of the claims in this specification (currently or subsequently appended to this specification). It should be noted that in the event of any inconsistency or conflict between the descriptions, definitions, and / or terminology used in the supplementary materials to this specification and the content of this specification, the descriptions, definitions, and / or terminology used in this specification shall prevail.

[0108] Finally, it should be understood that the embodiments described in this specification are merely illustrative of the principles of the embodiments described herein. Other variations may also fall within the scope of this specification. Therefore, alternative configurations of the embodiments described herein are intended to be illustrative rather than limiting, and should be considered consistent with the teachings of this specification. Accordingly, the embodiments described herein are not limited to those explicitly introduced and described herein.

Claims

1. A multi-user screen mirroring method, the method comprising: Multiple screen projection signals are acquired, and each of the multiple screen projection signals corresponds to the user identity of each of the multiple users; Obtain user voice information of at least one of the plurality of users; The method of casting the multiple screen projection signals is determined based on the user's voice information, and the method of casting the multiple screen projection signals includes switching the multiple screen projection signals; The method for determining the projection method of the plurality of projection signals includes: The user identity corresponding to the user voice information is determined based on the user voice information using a feature extraction model. The feature extraction model is trained on a consistency model, which is a machine learning model. The consistency model includes a first voice feature extraction layer, a second voice feature extraction layer, and a judgment layer. The first voice feature extraction layer processes the first user voice information to determine a first voice feature vector. The second voice feature extraction layer processes the second user voice information to determine a second voice feature vector. The judgment layer processes the first voice feature vector and the second voice feature vector to determine whether the first user voice information and the second user voice information belong to the same user. The target projection signal is determined based on the correspondence between the at least one user identity and the plurality of projection signals; Based on the target projection signal, a switching command for the projection signal of the projection device is determined.

2. The multi-user screen casting method according to claim 1, wherein determining the casting method for the plurality of screen casting signals based on the user voice information further includes: Determine the user identity of the current speaker based on the user's voice information; Convert the user's voice information into text information; Based on the text information, the probability that the text information is a declarative sentence is determined by a text judgment model. When the possibility meets the preset conditions, the screen projection device is controlled to switch the screen projection signal corresponding to the user identity.

3. The multi-user screen casting method according to claim 1, further comprising: Get the keyboard and mouse actions of the current speaker; Determine whether the keyboard and mouse actions are applied to the screen mirroring application; In response to the keyboard and mouse actions applied to the screen projection application, the projection method of the multiple screen projection signals is determined.

4. The multi-user screen casting method according to claim 1, wherein the feature extraction model includes a third speech feature extraction layer and a feature matching layer; the third speech feature extraction layer is determined based on the first speech feature extraction layer or the second speech feature extraction layer of the consistency model.

5. A multi-user screen mirroring system, the system comprising: The first acquisition module is used to acquire multiple screen projection signals, and each of the multiple screen projection signals has a corresponding relationship with the user identity of each of the multiple users; The second acquisition module is used to acquire user voice information of at least one of the plurality of users; The first determining module is used to determine the projection method of the plurality of projection signals based on the user voice information, wherein the projection method includes switching the plurality of projection signals; The first determining module is further configured to: determine the at least one user identity corresponding to the user voice information based on the user voice information using a feature extraction model, wherein the feature extraction model is trained based on a consistency model, the consistency model being a machine learning model, and the consistency model including a first voice feature extraction layer, a second voice feature extraction layer, and a judgment layer; the first voice feature extraction layer is configured to process the first user voice information to determine a first voice feature vector; the second voice feature extraction layer is configured to process the second user voice information to determine a second voice feature vector; the judgment layer is configured to process the first voice feature vector and the second voice feature vector to determine whether the first user voice information and the second user voice information belong to the same user; determine a target projection signal based on the correspondence between the at least one user identity and the plurality of projection signals; and determine a switching instruction for the projection signal of the projection device based on the target projection signal.

6. The multi-user screen projection system according to claim 5, wherein the first determining module is further configured to: Determine the user identity of the current speaker based on the user's voice information; Convert the user's voice information into text information; Based on the text information, the probability that the text information is a declarative sentence is determined by a text judgment model. When the possibility meets the preset conditions, the screen projection device is controlled to switch the screen projection signal corresponding to the user identity.

7. The multi-user screen projection system according to claim 5, further comprising: The third acquisition module is used to acquire the keyboard and mouse actions of the current speaker; The judgment module is used to determine whether the keyboard and mouse actions are applied to the screen mirroring application; The second determining module is used to determine the projection method of the plurality of projection signals in response to the keyboard and mouse actions acting on the projection application.

8. A multi-user screen projection device, characterized in that, The device includes: At least one storage medium that stores computer instructions; At least one processor executes the computer instructions to implement the multi-user screen projection method according to any one of claims 1 to 4.

9. A computer-readable storage medium, characterized in that, The storage medium stores computer instructions. When the computer reads the computer instructions, the computer executes the multi-user screen projection method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Automatic switching method for Microsoft voice recognition configuration files and system of Microsoft voice recognition configuration files

    CN104021146A

  • Identity identification and voice interaction operating method and device

    CN105895096A

  • Screen projection system and control method

    CN109753259A