Voice interaction method based on large language model and related device

By using a speech interaction method based on a large language model in the conference environment, combining real-time audio streaming and voice interaction interface, the efficiency of interactive behavior tracking and review in the conference is solved, achieving more intuitive and convenient understanding of interactive states, and improving user experience.

CN120091103APending Publication Date: 2025-06-03BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510272830.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

During the meeting, how to track and review the communication and interaction behaviors that occur in the meeting more efficiently and experiencingly, and improve the quality and efficiency of tasks and matters.

Method used

Using a speech interaction method based on a large language model, a real-time audio stream is collected in a physical environment, a user location is determined, and a user indicator is presented in the voice interaction interface, and its visual presentation attributes are adjusted to reflect the user's interaction status.

Benefits of technology

This method enables users to understand the interaction status and interaction between users in the meeting more intuitively and conveniently, reducing the interaction complexity and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091103A_ABST
    Figure CN120091103A_ABST
Patent Text Reader

Abstract

The invention provides a voice interaction method based on a large language model and a related device, and relates to the technical field of artificial intelligence such as voice recognition, audio processing, computer vision and large language models. The method comprises the following steps: determining a user included in a physical environment and a first position of the user in the physical environment based on a real-time audio stream collected in the physical environment; in a voice interaction interface presented for the physical environment, a user indicator corresponding to the user is presented in a manner of being associated with the target indicator, and the relative position relationship between the user indicator and the target indicator is determined based on the relative position relationship between the first position and a second position corresponding to the target indicator in the physical environment; a visual presentation attribute of the user indicator is adjusted based on a portion of the real-time audio stream corresponding to the user. Therefore, the user can more intuitively and conveniently understand the interaction state and the interaction condition between the users in the conference, the interaction complexity of the user is reduced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, specifically to artificial intelligence technology fields such as speech recognition, audio processing, computer vision, large language models, etc., and particularly to a speech interaction method, device, electronic device, computer-readable storage medium, and computer program product based on a large language model. Background Art

[0002] In work and life, when people handle complex tasks or matters that require multi-person collaboration, they usually adopt the form of meetings to communicate and discuss tasks and matters. Correspondingly, the centralized discussion achieved through meetings can improve the quality and efficiency of handling tasks and matters.

[0003] In such a background, how to help people conduct meetings more efficiently and with better experience, and facilitate people to track and review the communication and interaction behaviors that occur during meetings, is worthy of attention and an urgent need. Summary of the Invention

[0004] Embodiments of the present disclosure propose a speech interaction method, device, electronic device, computer-readable storage medium, and computer program product based on a large language model.

[0005] In a first aspect, embodiments of the present disclosure propose a speech interaction method based on a large language model, including: determining users included in a physical environment and a first position where the users are located in the physical environment based on a real-time audio stream collected in the physical environment; presenting, in a speech interaction interface presented for the physical environment, a user indicator corresponding to the users in association with a target indicator, where the relative position relationship between the user indicator and the target indicator is determined based on the relative position relationship between the first position and a second position corresponding to the target indicator in the physical environment; and adjusting a visual presentation attribute of the user indicator based on a part of the real-time audio stream corresponding to the users.

[0006] In a second aspect, embodiments of the present disclosure propose a speech interaction device based on a large language model, including: a user recognition and positioning unit configured to determine users included in a physical environment and a first position where the users are located in the physical environment based on a real-time audio stream collected in the physical environment; a user indicator presentation unit configured to present, in a speech interaction interface presented for the physical environment, a user indicator corresponding to the users in association with a target indicator, where the relative position relationship between the user indicator and the target indicator is determined based on the relative position relationship between the first position and a second position corresponding to the target indicator in the physical environment; and an indicator adjustment unit configured to adjust a visual presentation attribute of the user indicator based on a part of the real-time audio stream corresponding to the users.

[0007] In a third aspect, embodiments of the present disclosure provide an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to implement the large language model-based voice interaction method described in any implementation manner of the first aspect.

[0008] In a fourth aspect, embodiments of the present disclosure provide a non-transitory computer-readable storage medium storing computer instructions, and when the computer instructions are executed by a computer, the computer is enabled to implement the large language model-based voice interaction method described in any implementation manner of the first aspect.

[0009] In a fifth aspect, embodiments of the present disclosure provide a computer program product including a computer program, and when the computer program is executed by a processor, the computer program is enabled to implement the large language model-based voice interaction method described in any implementation manner of the first aspect.

[0010] For the large language model-based voice interaction method, apparatus, electronic device, computer-readable storage medium, and computer program product provided by the embodiments of the present disclosure, first, based on the real-time audio stream collected in the physical environment, the users included in the physical environment and the first positions of the users in the physical environment are determined; then, in the voice interaction interface presented for the physical environment, a user indicator corresponding to the user is presented in association with a target indicator, wherein the relative position relationship between the user indicator and the target indicator is determined based on the relative position relationship between the first position and the second position corresponding to the target indicator in the physical environment; finally, based on the part corresponding to the user in the real-time audio stream, the visual presentation attributes of the user indicator are adjusted.

[0011] The present disclosure enables users to more intuitively and conveniently understand the interaction status and interaction situation among users in a meeting, reduces the interaction complexity of users, and improves the user experience.

[0012] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Other features, objects, and advantages of the present disclosure will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings: Figure 1 is an exemplary system architecture to which the present disclosure can be applied; Figure 2A flowchart of a voice interaction process based on a large language model provided by an embodiment of the present disclosure; Figure 3 A flowchart of a process for determining identity information corresponding to a user provided by an embodiment of the present disclosure; Figures 4a - 4h Schematic diagrams of the effects of a voice interaction interface provided by embodiments of the present disclosure respectively; Figure 5 A schematic diagram of the effects achieved by a voice interaction process based on a large language model in an application scenario provided by an embodiment of the present disclosure; Figure 6 A structural block diagram of a voice interaction device based on a large language model provided by an embodiment of the present disclosure; Figure 7 A structural schematic diagram of an electronic device suitable for executing a voice interaction method based on a large language model provided by an embodiment of the present disclosure. Detailed implementation manners

[0014] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below. It should be noted that, without conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0015] In addition, in the technical solutions involved in the present disclosure, the acquisition, storage, use, processing, transportation, provision, and disclosure of user personal information (such as the real-time audio stream involved in the subsequent content of the present disclosure) and other processes all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0016] Figure 1 An exemplary system architecture 100 showing embodiments to which the voice interaction method, device, electronic device, and computer-readable storage medium based on a large language model of the present disclosure can be applied is shown.

[0017] As Figure 1As shown, the system architecture 100 may at least include terminal devices 101 and 102. For example, the terminal devices 101 and 102 may be arranged in a meeting environment or may be used in a meeting environment. For example, they are hardware devices such as smart screens, tablets, and laptop computers with computing and processing capabilities. In some scenarios, the terminal devices 101 and 102 may also be software. When the terminal devices 101 and 102 are software, they can be installed in the above-listed electronic devices and can be implemented as multiple software or software modules, or can be implemented as a single software or software module, which is not specifically limited here.

[0018] In some embodiments, the system architecture 100 may further include a network 103 and a server 104. The network 103 is used to provide a medium for the communication link between the terminal devices 101 and 102 and the server 104. The network 103 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0019] Similarly, the server 104 may also be hardware or software. When the server 104 is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server; when the server is software, it can be implemented as multiple software or software modules, or can be implemented as a single software or software module, which is not specifically limited here.

[0020] In such a case, in the system architecture 100, the server 104 can actually provide computing power for the terminal devices 101 and 102, and use the terminal devices 101 and 102 as the "presentation end" to present the processing results of the server 104.

[0021] For example, in the above meeting scenario, in the system architecture 100, the terminal devices 101 and 102 can actually be used as the "presentation terminals" to provide voice interaction interfaces 107 and 108 for users 105 and 106, so that the users 105 and 106 can use the voice interaction interfaces 107 and 108 to obtain the processing results of the server 104.

[0022] For ease of understanding, the system architecture 100 including the server 104 is used as an example. In such a case, the users 105 and 106 can use the terminal devices 101 and 102 to interact with the server 104 through the network 103 to receive or send messages, etc. Various applications for realizing information communication between the two can be installed on the terminal devices 101 and 102 and the server 104, such as meeting assistant applications, meeting record applications, instant messaging applications, etc.

[0023] Accordingly, the server 104 can also provide various services through various built-in applications. Taking the meeting assistant application that can provide meeting process records and presentations as an example, when the server 104 runs this meeting assistant application, the following effects can be achieved: First, the server 104 obtains the real-time audio stream collected from the physical environment from the terminal devices 101 and 102 through the network 103, and determines the users included in the physical environment and the first positions of the users in the physical environment; then, in the voice interaction interface presented for the physical environment, the server 104 presents user indicators corresponding to the users in association with the target indicator, wherein the relative position relationship between the user indicator and the target indicator is determined based on the relative position relationship between the first position and the second position corresponding to the target indicator in the physical environment; finally, the server 104 adjusts the visual presentation attributes of the user indicator based on the part of the real-time audio stream corresponding to the user.

[0024] Since operations such as determining the users included in the physical environment and the first positions of the users in the physical environment may require a large amount of computing resources and strong computing capabilities, generally, the speech interaction method based on the large language model provided in the subsequent embodiments of the present disclosure is executed by the server 104 with strong computing capabilities and a large amount of computing resources. Accordingly, the speech interaction device based on the large language model is generally also set in the server 104. However, it should also be pointed out that when the terminal devices 101 and 102 also have computing capabilities and computing resources that meet the requirements, the terminal devices 101 and 102 can also complete the above-mentioned various operations originally performed by the server 104 through the meeting assistant applications installed thereon, and then output the same results as the server 104. Especially in the case where there are multiple terminal devices with different computing capabilities at the same time, when the meeting assistant application determines that the terminal device where it is located has strong computing capabilities and a large amount of remaining computing resources, the terminal device can be allowed to perform the above operations, thereby appropriately reducing the computing pressure on the server 104. Accordingly, the speech interaction device based on the large language model can also be set in the terminal devices 101 and 102. For example, considering requirements such as convenience, when the computing capabilities and computing resources provided by the terminal devices 101 and 102 can also meet the requirements of the above-mentioned various operations, the terminal devices 101 and 102 can directly complete the above-mentioned various operations originally performed by the server 104. Accordingly, the exemplary system architecture 100 may not include the network 103 and the server 104.

[0025] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0026] Please refer to Figure 2 forFigure 2 FIG. 200 is a flowchart of a speech interaction process based on a large language model provided by an embodiment of the present disclosure.

[0027] Process 200 specifically includes the following steps: Step 201: Based on the real-time audio stream collected in the physical environment, determine the users included in the physical environment and the first position where the users are located in the physical environment; In the embodiments of the present disclosure, for the convenience of understanding, the execution subject of the speech interaction method based on the large language model can be directly implemented by, for example, Figure 1 the terminal devices 101 and 102 shown. For example, the terminal devices 101 and 102 can be "intelligent screens" arranged in a meeting environment with sufficient computing power and computing resources.

[0028] In this step, the execution subject can determine the users included in the physical environment and the position where the users are located in the physical environment (for the convenience of description, it is referred to as the "first position") based on the real-time audio stream collected in the physical environment. For example, the physical environment can be a real "meeting environment", and in such a meeting environment, users such as 105 and 106 can achieve a "meeting" through communication and interaction among users.

[0029] In this embodiment, the real-time audio stream can be obtained by a sound collection device such as a microphone configured by the execution subject through sound collection of the physical environment.

[0030] In some embodiments, the real-time audio stream can also be collected by a sound collection device arranged in the physical environment that is independent of the configuration of the execution subject but can communicate with the execution subject (for example, a microphone set independently of the "intelligent screen", etc.). Correspondingly, after collecting the real-time audio stream, the execution subject can determine the users included in the physical environment and the first position where the users are located in the physical environment through parsing the real-time audio stream.

[0031] Generally, the execution subject can at least parse the "tone" and "timbre" included in the real-time audio stream through standards such as "tone" and "timbre", and then determine the users involved and included in the real-time audio stream according to different "tones" and "timbres".

[0032] In some embodiments, users can record their information by pre-storing their timbres in the execution subject in advance, so that the execution subject can more efficiently and accurately "identify users" through timbres in the future. Such a method can not only enable the execution subject to more accurately identify and distinguish users, but also save the computing power consumption of the execution subject when performing parsing operations.

[0033] During this process, the executing entity can also determine the first position of the user in the physical environment by means of sound source localization and through parsing the real-time audio stream.

[0034] It should be understood that the acquisition of the real-time audio stream can be started according to the user's action instructions (for example, clicking on a specific control or making a specific gesture). Correspondingly, if no sound signal is detected in the physical environment, the executing entity can first choose to enter a waiting state and continuously perform detection until a sound signal is found.

[0035] Step 202: Present a user indicator corresponding to the user in association with a target indicator in the voice interaction interface presented for the physical environment; In the embodiments of the present disclosure, the executing entity can present a voice interaction interface for the physical environment. For example, when an application such as a meeting assistant is opened, the executing entity can first pre-generate and present such a voice interaction interface to feedback information to the user through this voice interaction interface. And because such a voice interaction interface is "automatically" generated by the executing entity, this way can also reduce the content that the user needs to configure in advance, reduce the interaction complexity of the user, and improve the user experience.

[0036] In the voice interaction interface, the position of the user in the physical environment and the relative positions between users can be presented through the user indicator corresponding to the user. And because such a user indicator corresponds to the user (or rather, the user often can have a "user indicator" uniquely corresponding to himself / herself), the interaction state of the user can also be presented through the user indicator. For example, by adjusting the visual style of the user indicator and adding dynamic effects, etc., to feedback and present the speaking and interaction states of the user.

[0037] Correspondingly, in order to better present the layout situation between users, the voice interaction interface pre-generated and formed by the executing entity can also include a target indicator. This target indicator corresponds to a real position in the physical environment (for convenience of description, it can be called the "second position"). Thus, the executing entity uses the target indicator as an anchor point to determine the layout of the user indicators and present the layout between users. For example, the target indicator can be a circular icon used to mark the "second position" in the voice interaction interface.

[0038] Or rather, the target indicator can be used as a positioning reference to present the layout of the user indicators. That is, the executing entity can simulate and present the relative position relationship between the first position where the user is located in the physical environment and the second position based on the target indicator and the user indicator.

[0039] In some embodiments, the second position may actually be the position where a sound collection device is set in the physical environment (which can be either a sound collection device arranged inside the execution entity or a sound collection device independent of the execution entity). Thus, the user can understand the relative position relationship between himself and the sound collection device through the voice interaction interface and adaptively adjust the interaction strategy (for example, whether to approach or move away from the sound collection device, etc.).

[0040] Correspondingly, in this step, if the execution entity identifies the user based on the real-time audio stream (for the first time), the execution entity can present a user indicator corresponding to the user in the voice interaction interface presented for the physical environment, in association with the target indicator. As described above, the relative position relationship between the user indicator and the target indicator can be determined based on the relative position relationship between the first position and the second position corresponding to the target indicator in the physical environment.

[0041] In practice, for the added position of the user indicator in the voice interaction interface, in terms of direction, it can completely refer to and restore the relationship between the first position and the second position, and in terms of the actual distance from the target indicator, the distance between the user indicator and the target indicator in the voice interaction interface can be determined by scaling the ratio of the distance between the first position and the second position.

[0042] In some alternative implementation manners of this embodiment, a distance upper limit value can also be pre-configured for the distance between the user indicator and the target indicator presented in the voice interaction interface (that is, if the determined distance between the user indicator and the target indicator in the voice interaction interface exceeds this distance upper limit, the distance upper limit value is actually adopted as the distance), so as to avoid the overall layout being scattered due to the user indicator being too far away and affecting the viewing experience.

[0043] It should be understood that the visual styles of the user indicator and the target indicator can be either the same icon, such as both being "circular icons", or different icons, such as the target indicator being "circular" and the user indicator being "triangular", etc. Similarly, in some scenarios, the user indicator and the target indicator can also be distinguished by different colors, which will not be elaborated here.

[0044] Step 203: Adjust the visual presentation attributes of the user indicator based on the part corresponding to the user in the real-time audio stream.

[0045] In an embodiment of the present disclosure, after adding the user indicator based on the above step 202, the execution entity may adjust the visual presentation attributes of the user indicator based on the part corresponding to the user in the real-time audio stream (i.e., the real-time audio stream of the part belonging to the user in the real-time audio stream, or the sub-real-time audio stream), so as to respond and feedback to the user's meeting interaction state in a targeted and real-time manner through the changes in the visual attributes of the user indicator, such as visual style, visual elements, color, size, presentation position, etc.

[0046] For example, for the user's speaking behavior in the real-time audio stream, etc., the execution entity may choose to provide a new visual effect different from the original visual effect by rotating the user indicator, highlighting the user indicator, changing the color of the user indicator, etc., so that the user can understand the meeting state of the user by observing the changes in the visual attributes of the user indicator in the voice interaction interface. For example, when the user indicator is (continuously) rotated, it means that the corresponding user "is speaking".

[0047] The voice interaction method based on the large language model provided by the embodiments of the present disclosure first determines the users included in the physical environment and the first positions of the users in the physical environment based on the real-time audio stream collected in the physical environment; then, in the voice interaction interface presented for the physical environment, a user indicator corresponding to the user is presented in association with the target indicator, wherein the relative position relationship between the user indicator and the target indicator is determined based on the relative position relationship between the first position and the second position corresponding to the target indicator in the physical environment; finally, the visual presentation attributes of the user indicator are adjusted based on the part corresponding to the user in the real-time audio stream. Thereby, it is convenient for users to more intuitively and conveniently understand the interaction state and interaction situation among users in the meeting, reduces the interaction complexity of users, and improves the user experience.

[0048] In some embodiments, during the process of the execution entity determining the users included in the physical environment and the first positions of the users in the physical environment (for example, during the execution of the above step 201), it may choose to first use the Time Difference of Arrival (TDOA) algorithm to determine the source azimuth and distance of the voice signal in the real-time audio stream collected in the physical environment.

[0049] The Time Difference of Arrival (TDOA) algorithm is a technology that determines the position of the signal source by calculating the difference in the arrival time of the signal at different receiving stations. The TDOA algorithm can calculate the source position of the voice signal by measuring the time difference of the voice signal from the transmission source to different receiving points and using these time differences.

[0050] Then, the execution entity can determine the users included in the physical environment based on the source orientation (for example, determine each included user based on different orientations and different timbres). For example, the execution entity can calculate the source orientation and distance of the voice signal in the physical environment by comparing the arrival times of different sound source signals (i.e., voice signals) based on the time difference location algorithm.

[0051] Finally, the execution entity determines the first position of the user in the physical environment based on the distance. Thus, the execution entity can use the time difference location algorithm to more accurately locate, identify, and separate users.

[0052] In some embodiments, as discussed above, in a physical environment such as a meeting environment, there may be at least two or more users. In such a case, the execution entity can also respond by parsing the respective parts corresponding to each user in the real-time audio stream, for example, by means of Blind Source Separation (BSS), to split the "real-time audio stream" and determine the respective "parts" and "sub-real-time audio streams" corresponding to each user.

[0053] Blind source separation is a technique for recovering independent source signals from multiple mixed signals, which can be separated only based on the received mixed signals. Accordingly, using blind source separation enables the execution entity to independently process real-time audio streams involving multiple users, to solve the problem of voice mixing when multiple people speak simultaneously, and to ensure that the voice content of each user can be accurately recognized and personalized responses can be made.

[0054] For the part of the real-time audio stream corresponding to the user (or in the case of multiple users, it can also be for the part of the real-time audio stream corresponding to each user), the execution entity can choose to determine the text information of the content included in this part by parsing the (voice) content of this part (i.e., the "content" in text form). For example, based on Automatic Speech Recognition (ASR) technology, the part of the real-time audio stream corresponding to the user can be converted into corresponding text information.

[0055] Then, the execution entity can determine the identity information corresponding to each user based on the text information. This identity information can be directly specific to the user's name (for example, the user has provided the name to the execution entity in advance through authorization and input methods), or it can only be located to the user's role. For example, such identity information can be Speaker 1, Speaker 2, or Participant 1, Participant 2, etc. For example, the execution entity can determine the corresponding identity information (such as the host, the explainer, the questioner, etc.) based on the semantic information summarized from the text information.

[0056] Next, the execution entity can further present an identity prompt associated with the user identifier corresponding to the user based on the identity information corresponding to the user. For example, in terms of visual style, the identity prompt can be the text information corresponding to the identity information.

[0057] Thus, this enables the execution entity to not only recognize the content of the speech, but also identify the speaking user based on the spatial position of the speech signal, ensuring personalized responses in a multi-user environment. And through the identity prompt, users can more intuitively and effectively determine the participants in the meeting, as well as the identities and roles that the participants have in the meeting.

[0058] In some embodiments, to avoid incorrect user identification caused by inaccurate positioning, tone color errors, etc., for example, identifying the speech and interaction of one user as two or more users. The execution entity can also choose to determine whether there is a situation where the same user is determined to have at least two different identity information.

[0059] Correspondingly, if there is such a situation, the execution entity can respond to this by merging at least two different identity information (for example, adjusting the subsequent identity information to the first determined "identity information" to complete the merger), and merging the parts in the real-time audio stream used to determine at least two different identity information (for example, merging the various "parts" that were previously incorrectly split and extracted from each real-time audio stream into a whole). Thus, to avoid response chaos caused by incorrect user positioning and identification.

[0060] Regarding this, for ease of understanding, reference can be made to Figure 3 for illustration. Figure 3 FIG. is a flowchart of a process for determining the identity information corresponding to a user provided by an embodiment of the present disclosure, which includes process 300.

[0061] Process 300 specifically includes the following steps: Step 301: Parse the text information of the part corresponding to the user in the real-time audio stream; Step 302: Determine the identity information corresponding to the user based on the text information; As discussed above in steps 301 and 302, the execution entity can first parse the text information based on automatic speech recognition technology, and then determine the identity information corresponding to each user based on the identity information, which will not be repeated here.

[0062] Step 303: Combine the text information corresponding to the first user and the text information corresponding to the second user to obtain combined text information; Specifically, as discussed above, if the execution entity detects that there are at least two "users", the execution entity can select any two "users" as the first user and the second user, and combine the text information corresponding to the first user and the text information corresponding to the second user to obtain combined text information.

[0063] Step 304: Perform context analysis on the combined text information using a large language model to obtain a context analysis result; Specifically, based on the above step 303, the execution entity can use a large language model to perform context analysis on the combined text information to determine whether the combined information belongs to a "complete event" described and stated by the same person in terms of semantics analyzed through context, for example, and whether the combined information comes from the same person in terms of context, and accordingly obtain a context analysis result.

[0064] Large Language Model (LLM for short). An LLM is an artificial intelligence model designed to understand and generate human language, and an LLM can perform corresponding processing operations based on the content it understands to obtain corresponding processing results. For example, after obtaining the "combined text information", the LLM can, after understanding the instruction (such as identifying whether the combined text information comes from the same person), determine whether the combined text information comes from and belongs to the same person (or rather, user) based on its semantic understanding and processing of the combined text information.

[0065] Generally, an LLM can be trained on a large amount of text data and can perform a wide range of tasks, including text summarization, translation, sentiment analysis, and so on. The characteristic of an LLM is its large scale. Usually, it can include a large number of parameters to help them learn complex patterns in language data. These models are usually based on deep learning architectures such as transformers, which helps them provide better processing performance on various NLP tasks.

[0066] In addition, for the LLM model, it can omit the "prompt" in a default configuration manner. For example, after obtaining the "combined text information", for the purpose of determining whether the combined text information comes from the same person based on its semantic understanding and processing of the combined text information, the LLM can, based on the default configuration, naturally understand the operation that needs to be performed on the input "combined text information" (that is, determining whether the combined text information comes from the same person). Thus, in a default configuration manner, the generative model can stably and directionally process the combined text information and improve the efficiency of using the generative large language model.

[0067] Accordingly, the context analysis result can indicate whether the first user and the second user are the same user with different identity information. For example, if the combined information belongs to the "complete event" described and stated by the same person, the context analysis result can indicate that the first user and the second user are actually the "same user".

[0068] Step 305: In response to determining that the same user is determined to have at least two different identity information, merge at least two different identity information and the part of the real-time audio stream used to determine at least two different identity information.

[0069] Specifically, as discussed above, in this step, the execution entity can perform the "merge" action when it is determined that the same user is determined to have at least two different identity information (for example, based on the context analysis result obtained in step 304, it is determined that the same user is determined to have at least two different identity information), and the description will not be repeated here.

[0070] Thus, on the basis of integrating context information, the large language model is also used to perform deeper understanding and dialogue management.

[0071] It should be understood that in the embodiments using the large language model, the execution entity can also choose to hand over one or more tasks among parsing text information, determining identity information, and combining text information to the large model for processing. Thus, while making more full use of the computing power of the large language model, it is also possible to enable the processing parameters of these tasks to be learned by the large language model and used as a reference for other tasks, so that the large language model can complete the processing of these tasks more uniformly and reduce the result and style fragmentation that may occur due to model crossing.

[0072] In some embodiments, in order to make more full use of the computing power of the large language model, the execution entity can also hand over the step of parsing the real-time audio stream as discussed above to the large language model for processing. For example, in the process of executing the above step 201, the execution entity can actually choose to call the large language model to determine the users included in the physical environment and the first position of the users in the physical environment based on the real-time audio stream collected in the physical environment. For example, the large language model can also use the time difference positioning algorithm or other algorithms to parse the real-time audio stream to determine the users included in the physical environment and the first position of the users in the physical environment.

[0073] For another example, the process of "adjusting the visual presentation attributes of the user indicator" to be discussed and described later can also be implemented by the large language model (for example, using the large language model to generate a dynamic user indicator, generating a user indicator with a changed size to achieve the adjustment of "size", etc.).

[0074] It should be understood that after completing one round of user merging, the merged user can still be determined in a similar manner as to whether it can be further merged with other "users". Thus, through such "continuous merging", the parts belonging to the same user in the real-time audio stream can be continuously integrated. Through such continuous integration, the executing entity can ensure the continuity and fluency of information exchange in the conversation, and avoid situations such as information chaos, interruption, or restart.

[0075] In some embodiments, the executing entity can also choose to continuously track and store the text information of each user and the corresponding part of the real-time audio stream. Thus, by retaining the conversation context and interaction records of each user, the executing entity can better understand the previous inputs and needs of the user, and thus better understand and respond in subsequent interactions (for example, in some embodiments, the executing entity can also provide personalized response services such as online Q&A and task processing for the user through, for example, large language models), and further avoid situations such as discontinuous conversations and information confusion.

[0076] In some alternative implementation manners of this embodiment, if the executing entity is configured to select and determine the identity information of the user, then in such a case, for the above-discussed "target indicator", a target user with a specific identity can also be selected as the "anchor point". For example, the position of the "target indicator" in the physical environment can actually be the position corresponding to the target user with identity information such as "presenter" or "host". That is, the above-mentioned second position can be, in addition to the position where the sound collection device is set in the physical environment, also the position where the target user is located in the physical environment.

[0077] Thus, the executing entity can use the position of the target user with specific identity information in the physical environment as a reference to present the environment (or rather, the meeting layout), so that the executing entity can compose the voice interaction interface based on the interaction mode between users (for example, the mode of presenter and audience), highlighting the "interaction sense" corresponding to the interaction mode between users.

[0078] In some embodiments, for the above-mentioned target indicator, its presentation position can be located at the center of the voice interaction interface. That is, the executing entity can choose to present the target indicator at the center of the voice interaction interface.

[0079] Thus, the user layout centered around the target indicator in the voice interaction interface can be entirely located at the center of the voice interaction interface, so as to avoid excessive deviation of the user indicator and affect the viewing experience of the user, thereby enhancing the user experience.

[0080] In some embodiments, in order to give users a more "immersive" experience, when forming a voice interaction interface, the execution entity may also refer to the physical environment to enhance the user experience.

[0081] Specifically, during the process of forming a voice interaction interface, the execution entity may select a planar layout based on the physical environment to form the voice interaction interface. For example, the execution entity may determine the planar layout of the physical environment from the layout template and / or layout diagram of the physical environment selected and provided by the user, or the execution entity may also determine the three-dimensional layout of the physical environment through a collection device such as a camera, and determine the planar layout based on this three-dimensional layout. Then, based on this planar layout, the execution entity forms a voice interaction interface (for example, directly uses this planar layout or uses it after scaling down proportionally).

[0082] Thus, through the voice interaction interface constructed based on the actual situation of the physical environment, it can more realistically feedback the situation during the meeting process, facilitating users to more effectively read information such as positions, and enhancing the user experience.

[0083] In some alternative implementation manners of this embodiment, in such a case, for the "target indicator", the execution entity can determine the planar position of the second position in the physical environment corresponding to the target indicator in the planar layout. Then, through a mapping method, based on the planar position, determine the position to present the target indicator in the voice interaction interface and present it. Thus, the layout of the target indicator and the user indicator can be closer to the "actual situation" in the physical environment.

[0084] In some embodiments, if there is already a formed voice interaction interface before audio collection, the execution entity can also additionally select to present some "sound collection indicators" in such an interface. Correspondingly, after presenting the "sound collection indicators", during the process of performing real-time audio stream collection, the execution entity can also feedback the real-time audio stream collection process to the user by changing the visual style of the "sound collection indicators" (for example, "geometric small balls" that adaptively change based on timbre, pitch, speech rate, etc.).

[0085] Next, multiple embodiments will be further used to discuss the adjustment method of the visual presentation attributes of the user indicator by the execution entity.

[0086] In some embodiments, when the execution entity adjusts the visual presentation attributes of the user indicator, it can select to continuously adjust the size of the user indicator based on the cumulative number of words in the text information corresponding to the user in the real-time audio stream. The size is positively correlated with the cumulative number of words.

[0087] That is, the execution entity can adjust the size of the user indicator (in its visual presentation attributes) with reference to the "cumulative amount of speech content" of the user. For example, as the cumulative amount of words output by the user increases, the size of the user indicator gradually becomes larger correspondingly. Thus, it is possible to intuitively understand the cumulative amount of words output by the user during the interaction through the size of the user indicator.

[0088] For ease of understanding, reference can also be made to Figures 4a - 4h to illustrate the process and effects of the adjustment methods involved in each embodiment. Figures 4a - 4h They are all schematic diagrams of the effects of the voice interaction interface provided by the embodiments of the present disclosure.

[0089] In Figure 4a It includes a voice interaction interface 410 (for example, it can actually be the voice interaction interface 107 provided by the terminal device 101 and / or the voice interaction interface 108 provided by the terminal device 102). The voice interaction interface 410 includes a target indicator 411, a user indicator 412 corresponding to the user 105, and a user indicator 413 corresponding to the user 106. For ease of understanding, Figure 4a it can be simply understood as the "initial state" where the user indicators 412 and 413 have just been added and no adjustment has been made to the "user indicator 412" and "user indicator 413".

[0090] Exemplarily, the "user indicator 412" can have a larger size because the user 105 is closer to the "second position" corresponding to the target indicator 411 in the physical environment compared to the user 106.

[0091] Next, reference can be made to Figure 4b In Figure 4b the user indicator 412 can be adjusted by the execution entity to the user indicator 422 as the user 105's speech increases and the cumulative amount of words rises. The size of the user indicator 422 is larger than that of the user indicator 412.

[0092] In some alternative implementation manners of this embodiment, during the process of adjusting the size based on the cumulative amount of words, an upper limit size can also be set to prevent the user indicator from expanding indefinitely and causing chaos in the layout of the voice interaction interface. For example, the execution entity can set a cumulative amount threshold corresponding to the upper limit size, so that when the execution entity determines that the cumulative amount of words reaches this cumulative amount threshold, it will no longer continue to "expand" or "enlarge" the size.

[0093] That is, the execution entity can respond when it determines that the cumulative amount of words is greater than or equal to this cumulative amount threshold and stop continuously adjusting the size of the user indicator.

[0094] In some embodiments, the executing entity may also adjust the size of the user indicator based on the volume of the part corresponding to the user in the real-time audio stream. The size is positively correlated with the volume. For example, the executing entity may determine the actual enlarged size based on the product of the volume and a set proportionality coefficient.

[0095] In some alternative implementation manners of this embodiment, the executing entity may also determine a magnification factor based on the volume to adjust the size in a "doubling" or "shrinking" manner. Thus, the problem of inconsistent magnification logic caused by differences in the original sizes (for example, because the original sizes are inconsistent, when increasing the same fixed size, the user indicator with the smaller original size may be "overly enlarged") can be avoided.

[0096] Thus, the size of the user indicator can be used to intuitively reflect the speaking volume of the user.

[0097] Next, reference can be made to Figure 4c . In Figure 4c , the executing entity may adjust the user indicator 412 and the user indicator 413 due to the speaking behaviors of users 105 and 106 (for example, the user indicator 412 is adjusted to the user indicator 432, and the user indicator 413 is adjusted to the user indicator 433). During this process, exemplarily, the executing entity may further make the adjustment amplitude (for example, the magnification factor) of the user indicator 413 greater than that of the user indicator 412 because the speaking volume of user 106 is greater than that of user 105.

[0098] In addition, the executing entity may also adjust the position in the (historical, current) visual presentation attributes of the corresponding user indicator based on the speaking behavior of the user, so as to present the interaction situation of the user's speaking, for example, through the movement of the user indicator. Accordingly, users can understand the meeting interaction status such as the cumulative number of words generated by the user due to speaking through the position change of the user indicator.

[0099] In some embodiments, when adjusting the visual presentation attributes of the user indicator, the executing entity may choose to continuously adjust the presentation position of the user indicator to move towards the target indicator based on the cumulative number of words in the text information of the part corresponding to the user in the real-time audio stream. The moving distance of the presentation position of the user indicator is positively correlated with the cumulative number of words.

[0100] That is, the user indicator can gradually and continuously move towards the target indicator as the cumulative number of words output by the user increases.

[0101] Next, reference can be made to Figure 4d . InFigure 4d In this case, as the user's speech increases, that is, as the cumulative word count rises, the user indicator 412 can be moved by the executing entity in the direction pointing to the target indicator 411, and the user indicator 412 is adjusted to the user indicator 442.

[0102] In some alternative implementation manners of this embodiment, during the process of moving the user indicator based on the cumulative word count, the presentation position of the target indicator can be selected as the upper limit position of the movement. That is, if the executing entity determines that such continuous movement makes the presentation position of the user indicator the same as the presentation position of the target indicator, the executing entity can respond thereto and stop continuously adjusting the presentation position of the user indicator to move towards the target indicator.

[0103] Thereby, the visual effect of the user indicator "gradually merging into and integrating with" the target indicator is achieved, so that the meeting participation, interaction status, and progress of the user can be intuitively presented through the positional relationship between the user indicator and the target indicator, enhancing the user experience.

[0104] For this, reference can be made to Figure 4e . In Figure 4e , after the user indicator 412 shown in Figure 4d moves in the direction pointing to the target indicator 411, when its presentation position reaches the target indicator 411 (for example, the geometric centers of the two overlap), the executing entity can stop further movement. Correspondingly, in such a case, the user indicator 412 can finally be adjusted to the user indicator 452 (exemplarily, the user indicator 452 becomes "not directly visible" because it has integrated into the target indicator 411).

[0105] In some alternative implementation manners of this embodiment, when allowing the user indicator to merge into and integrate with the target indicator, at each stage when the user indicator starts to integrate and is in the process of integrating, the executing entity can also present a fusion special effect based on the fusion position, so that the presented integration process is more "natural" and the user's viewing experience is enhanced.

[0106] In other embodiments, for the upper limit position of the process of moving the user indicator based on the cumulative word count, according to different requirements, in different embodiments, "there is an overlapping part between the user indicator and the target indicator" can also be selected as the upper limit position. That is, in other embodiments, the executing entity can also choose to respond to the existence of an overlapping part between the user indicator and the target indicator and stop continuously adjusting the presentation position of the user indicator to move towards the target indicator.

[0107] Thus, the "user indicator" can be presented more independently and clearly during movement to meet different users' different understandings and requirements for "clarity" and "intuition".

[0108] For this, reference can be made to Figure 4f . In Figure 4f , after the user indicator 412 shown in Figure 4d moves in the direction pointing to the target indicator 411, in the case where the user indicator 412 collides with or overlaps the target indicator 411, the execution entity can stop moving further. Correspondingly, in such a case, the user indicator 412 can finally be adjusted to the user indicator 462.

[0109] Similarly, the execution entity can also choose to stop moving further, for example, just after the user indicator has moved to a state of being "completely merged" by the target indicator.

[0110] In some embodiments, the execution entity can also, as discussed above, reflect the state of the user, such as when the user is speaking, by directly changing the visual style of the user indicator.

[0111] Specifically, for the process of adjusting the visual presentation attributes of the user indicator based on the part corresponding to the user in the real-time audio stream, the execution entity can choose to adjust the visual style (in the visual presentation attributes) of the user indicator to a dynamic icon in response to determining that the user is currently speaking based on the part corresponding to the user in the real-time audio stream.

[0112] For example, the execution entity can form a dynamic icon by adding dynamic elements to the visual style of the user indicator or by moving the edge shape of the visual style of the user indicator according to a preset animation rule, and present the dynamic icon.

[0113] For example, in the case of a user indicator with a cartoon image as the visual style, the execution entity can present a "short video" formed based on the cartoon image as a dynamic icon. Thus, through the dynamic icon, the state that the user is speaking can be intuitively feedback.

[0114] In some embodiments, reference can also be made to the methods of size adjustment and position adjustment discussed above, and when it is determined that the user is speaking, the user indicator can be made "dynamic" in the form of cyclic "size adjustment" and "position movement" (for example, in such a "dynamic style", each cycle can start from the original position of the user indicator or the position adjusted based on the word count accumulation amount, and end at the presentation position of the target indicator).

[0115] In some embodiments, if the execution entity determines that the user is currently speaking, that is, the execution entity determines that the user is currently speaking based on the part corresponding to the user in the real-time audio stream, the execution entity can also respond thereto and select to present, in the voice interaction interface, between the user indicator and the target indicator, a dynamic indicator that starts from the presentation position of the user indicator and points to the presentation position of the target indicator.

[0116] For example, the dynamic indicator can be a continuously moving flow arrow. Thus, through such a dynamic indicator, the interactive state in which the user is speaking can be more intuitively reflected.

[0117] In some alternative implementation manners of this embodiment, the dynamic indicator can be a dynamic text stream generated based on the text information of the content that the user is currently speaking. Thus, through the dynamic text stream, the content that the user is currently telling can also be prompted, so as to provide auxiliary functions such as facilitating reading and understanding the speech content while enhancing the sense of interaction, and reducing the operation complexity of the user (for example, the user can directly obtain the speech content of other users through the dynamic text stream without having to additionally call function plugins such as subtitle adding plugins).

[0118] For this, reference can be made to Figure 4g . In Figure 4g , the execution entity can generate a dynamic text stream 414 according to the speech content "XXXXX" of the user 105 as the user indicator 412 "is speaking". Then, in the voice interaction interface 410, the execution entity can select to present, between the user indicator 412 and the target indicator 411, a dynamic text stream 414 that starts from the presentation position of the user indicator 412 and points to the presentation position of the target indicator 411.

[0119] In some embodiments, the execution entity can also add a text box corresponding to the user in the voice interaction interface (usually, the presentation position of the text box can be determined based on a pre-configured interface layout. For example, the presentation position of the text box can be the two side edges of the voice interaction interface). The text box is used to present the text information of the part corresponding to the user in the real-time audio stream. For example, the execution entity can provide a function control in the voice interaction interface so that the user can indicate the execution entity to add a text box by triggering the function control.

[0120] The text box is used to present the text information of the part corresponding to the user in the real-time audio stream. That is, the execution entity can set a corresponding text box for each user to present the text information of the content output by the user. Thus, it is convenient for the user himself or to retrospect and assist in understanding the content generated by other users' speeches.

[0121] For this, reference can be made toFigure 4h In Figure 4h , the execution entity (e.g., in response to the above function control being triggered) can add a text box 415 corresponding to user 105 and a text box 416 corresponding to user 106 in the voice interaction interface 410, so as to present the text content generated by user 105's speech through text box 415, and present the text content generated by user 106's speech through text box 416.

[0122] In practice, the text box can also represent its corresponding relationship with the user by presenting the identity information of user 105 and user 106.

[0123] In summary, the present disclosure can imitate the physical real world through the expression of the voice interaction interface. It can not only express the user's orientation information by mapping the three-dimensional space to the two-dimensional user interface, but also present the transmission process of the physical world sound in a graphical and dynamic way on the user interface, making it convenient for users to more intuitively and conveniently understand the interaction status and interaction situation among users in the meeting, reducing the user's interaction complexity and improving the user experience.

[0124] To deepen the understanding, the present disclosure also combines a specific application scenario and gives a specific implementation solution. Please refer to the Figure 5 effect schematic diagram of the effect achieved by the voice interaction process based on the large language model as shown. For easy understanding, it can be combined with the Figure 1 system architecture 100 shown for description.

[0125] In Figure 5 , the sound collection device 510 can collect the interaction behaviors of user 105 and user 106 (i.e., the speech interaction behaviors of providing voice). Then, the terminal device 101, which is exemplified as the execution entity, can obtain the real-time audio stream 515 collected by the sound collection device 510, and determine user 105 and user 106 and the first positions of user 105 and user 106 in the physical environment through the analysis of the real-time audio stream 515.

[0126] Next, the terminal device 101 presents the user indicator 522 corresponding to user 105 and the user indicator 523 corresponding to user 106 in the voice interaction interface 107 presented for the physical environment, in association with the target indicator 521.

[0127] Then, the terminal device 101 can adjust the visual presentation attributes of the user indicator 522 based on the part corresponding to the user 105 in the real-time audio stream 515 (e.g., move in the direction of the target indicator 521), and adjust the visual presentation attributes of the user indicator 523 based on the part corresponding to the user 106 in the real-time audio stream 515 (e.g., move in the direction of the target indicator 521).

[0128] Further reference Figure 6 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a speech interaction device based on a large language model. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0129] As Figure 6 shown, the speech interaction device 600 based on a large language model in this embodiment may include: a user identification and positioning unit 601, a user indicator presentation unit 602, and an indicator adjustment unit 603. Among them, the user identification and positioning unit 601 is configured to determine the users included in the physical environment and the first positions of the users in the physical environment based on the real-time audio stream collected in the physical environment; the user indicator presentation unit 602 is configured to present, in the speech interaction interface presented for the physical environment, the user indicators corresponding to the users in association with the target indicator, where the relative position relationship between the user indicator and the target indicator is determined based on the relative position relationship between the first position and the second position corresponding to the target indicator in the physical environment; the indicator adjustment unit 603 is configured to adjust the visual presentation attributes of the user indicator based on the part corresponding to the user in the real-time audio stream.

[0130] In this embodiment, in the speech interaction device 600 based on a large language model: the specific processing of the user identification and positioning unit 601, the user indicator presentation unit 602, and the indicator adjustment unit 603 and the technical effects brought by them can be respectively referred to Figure 2 the relevant descriptions of steps 201-203 in the corresponding embodiments, which will not be elaborated here.

[0131] In some optional implementation manners of this embodiment, the user identification and positioning unit 601 is further configured to use a time difference positioning algorithm to determine the source azimuth and distance of the speech signal in the real-time audio stream collected in the physical environment; determine the users included in the physical environment based on the source azimuth; and determine the first positions of the users in the physical environment based on the distance.

[0132] In some alternative implementation manners of this embodiment, the apparatus 600 further includes: a text information parsing unit configured to parse the text information corresponding to the user in the real-time audio stream; an identity information determination unit configured to determine the identity information corresponding to the user based on the text information; and an identity prompt presenting unit configured to present an identity prompt in association with the user indicator corresponding to the user based on the identity information corresponding to the user.

[0133] In some alternative implementation manners of this embodiment, the apparatus 600 further includes: a user merging unit configured to, in response to determining that the same user is determined to have at least two different identity information, merge the at least two different identity information and the part of the real-time audio stream used to determine the at least two different identity information.

[0134] In some alternative implementation manners of this embodiment, the apparatus 600 further includes: a text information combining unit configured to combine the text information corresponding to a first user and the text information corresponding to a second user to obtain combined text information, where the identity information corresponding to the first user is different from the identity information corresponding to the second user; and a same user determination unit configured to perform context analysis on the combined text information by using a large language model to obtain a context analysis result, where the context analysis result indicates whether the first user and the second user are the same user with different identity information.

[0135] In some alternative implementation manners of this embodiment, the apparatus 600 further includes: a first target indicator presenting unit configured to present a target indicator at the center of the interface of the voice interaction interface.

[0136] In some alternative implementation manners of this embodiment, the apparatus 600 further includes: a voice interaction interface forming unit configured to form a voice interaction interface based on the planar layout of the physical environment.

[0137] In some alternative implementation manners of this embodiment, the apparatus 600 further includes: a second target indicator presenting unit configured to determine the planar position of a second position in the planar layout; and based on the planar position, present a target indicator in the voice interaction interface.

[0138] In some alternative implementation manners of this embodiment, the indicator adjustment unit 603 is further configured to continuously adjust the size of the user indicator based on the cumulative amount of the number of words of the text information corresponding to the user in the real-time audio stream, where the size is positively correlated with the cumulative amount of the number of words.

[0139] In some alternative implementation manners of this embodiment, the apparatus 600 further includes: a size adjustment stop unit configured to stop continuously adjusting the size of the user indicator in response to the cumulative amount of the number of words being greater than or equal to the cumulative amount threshold.

[0140] In some alternative implementation manners of this embodiment, the indicator adjustment unit 603 is further configured to adjust the size of the user indicator based on the volume of the part corresponding to the user in the real-time audio stream, where the size of the size is positively correlated with the volume of the volume.

[0141] In some alternative implementation manners of this embodiment, the indicator adjustment unit 603 is further configured to continuously adjust the presentation position of the user indicator to move towards the target indicator based on the cumulative number of words in the text information of the part corresponding to the user in the real-time audio stream, where the moving distance of the presentation position of the user indicator is positively correlated with the cumulative number of words.

[0142] In some alternative implementation manners of this embodiment, the apparatus 600 further includes: a first movement stop unit configured to stop continuously adjusting the presentation position of the user indicator to move towards the target indicator in response to the presentation position of the user indicator being the same as the presentation position of the target indicator.

[0143] In some alternative implementation manners of this embodiment, the apparatus 600 further includes: a second movement stop unit configured to stop continuously adjusting the presentation position of the user indicator to move towards the target indicator in response to there being an overlapping part between the user indicator and the target indicator.

[0144] In some alternative implementation manners of this embodiment, the indicator adjustment unit 603 is further configured to adjust the visual style of the user indicator to a dynamic icon in response to determining that the user is currently speaking based on the part corresponding to the user in the real-time audio stream.

[0145] In some alternative implementation manners of this embodiment, the apparatus 600 further includes: a dynamic indicator presentation unit configured to present a dynamic indicator starting from the presentation position of the user indicator and pointing to the presentation position of the target indicator between the user indicator and the target indicator in response to determining that the user is currently speaking based on the part corresponding to the user in the real-time audio stream.

[0146] In some alternative implementation manners of this embodiment, the dynamic indicator includes: a dynamic text stream generated based on the text information of the content that the user is currently speaking.

[0147] In some alternative implementation manners of this embodiment, the apparatus 600 further includes: a text box presentation unit configured to add a text box corresponding to the user in the voice interaction interface, where the text box is used to present the text information of the part corresponding to the user in the real-time audio stream.

[0148] In some alternative implementation manners of this embodiment, the second position includes: the position where a sound collection device is set in the physical environment, or the position where the target user is located in the physical environment.

[0149] This embodiment exists as a device embodiment corresponding to the above method embodiment. The speech interaction device based on the large language model provided in this embodiment determines the users included in the physical environment and the first position where the users are located in the physical environment based on the real-time audio stream collected in the physical environment; in the speech interaction interface presented for the physical environment, a user indicator corresponding to the user is presented in association with the target indicator, and the relative position relationship between the user indicator and the target indicator is determined based on the relative position relationship between the first position and the second position corresponding to the target indicator in the physical environment; based on the part corresponding to the user in the real-time audio stream, the visual presentation attributes of the user indicator are adjusted. Thus, it is convenient for users to more intuitively and conveniently understand the interaction status and interaction situation among users in the meeting, reduces the interaction complexity of users, and improves the user experience.

[0150] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0151] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0152] As Figure 7 shown, the device 700 includes a computing unit 701, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 702 or the computer program loaded from the storage unit 708 into the random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. The input / output (I / O) interface 705 is also connected to the bus 704.

[0153] Multiple components in device 700 are connected to I / O interface 705, including: input unit 706, such as a keyboard, mouse, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 708, such as a disk, optical disc, etc.; and communication unit 709, such as a network card, modem, wireless communication transceiver, etc. Communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0154] Computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of computing unit 701 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 701 executes the various methods and processes described above, such as the speech interaction method based on a large language model. For example, in some embodiments, the speech interaction method based on a large language model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by computing unit 701, one or more steps of the speech interaction method based on a large language model described above can be executed. Alternatively, in other embodiments, computing unit 701 can be configured to execute the speech interaction method based on a large language model in any other suitable way (e.g., by means of firmware).

[0155] The various embodiments of the systems and techniques described above in this article can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, which can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a dedicated or general-purpose programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0156] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0157] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0158] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0159] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0160] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to address the defects of high management difficulty and weak business scalability existing in traditional physical hosts and virtual private server (VPS) services. The server can also be a server of a distributed system, or a server combined with a blockchain.

[0161] According to the technical solution of the embodiment of the present disclosure, based on the real-time audio stream collected in the physical environment, determine the users included in the physical environment and the first positions of the users in the physical environment; in the voice interaction interface presented for the physical environment, present a user indicator corresponding to the user in association with a target indicator, and the relative position relationship between the user indicator and the target indicator is determined based on the relative position relationship between the first position and the second position corresponding to the target indicator in the physical environment; based on the part corresponding to the user in the real-time audio stream, adjust the visual presentation attributes of the user indicator. Thereby, it is possible to facilitate the user to more intuitively and conveniently understand the interaction status and interaction situation among users in the meeting, reduce the interaction complexity of the user, and improve the user experience.

[0162] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution provided by the present disclosure can be achieved, and no limitation is imposed herein.

[0163] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A voice interaction method based on a large language model, comprising: Determine, based on a real-time audio stream collected in a physical environment, a user included in the physical environment and a first position of the user in the physical environment; In a voice interaction interface presented for the physical environment, presenting a user indicator corresponding to the user in association with a target indicator, wherein a relative positional relationship between the user indicator and the target indicator is determined based on a relative positional relationship between the first position and a second position corresponding to the target indicator in the physical environment; A visual presentation attribute of the user indicator is adjusted based on the portion of the real-time audio stream corresponding to the user.

2. The method according to claim 1, wherein: The determining, based on the real-time audio stream collected in the physical environment, a user included in the physical environment and a first position of the user in the physical environment includes: Use the time difference positioning algorithm to determine the source direction and distance of the voice signal in the real-time audio stream collected in the physical environment; determining users included in the physical environment based on the source location; A first location of the user in the physical environment is determined based on the distance.

3. The method according to claim 1, further comprising: Parsing text information corresponding to the user in the real-time audio stream; Determine the identity information corresponding to the user based on the text information; Based on the identity information corresponding to the user, an identity prompt is presented in association with a user indicator corresponding to the user.

4. The method according to claim 3, further comprising: In response to determining that the same user is determined to have at least two different identity information, the at least two different identity information and the portion of the real-time audio stream used to determine the at least two different identity information are merged.

5. The method according to claim 4, further comprising: combining text information corresponding to a first user and text information corresponding to a second user to obtain combined text information, wherein identity information corresponding to the first user is different from identity information corresponding to the second user; A context analysis is performed on the combined text information using a large language model to obtain a context analysis result, wherein the context analysis result indicates whether the first user and the second user are the same user with different identity information.

6. The method according to claim 1, further comprising: The target indicator is presented at the center of the interface of the voice interaction interface.

7. The method according to claim 1, further comprising: The voice interaction interface is formed based on the planar layout of the physical environment.

8. The method according to claim 7, further comprising: determining a planar position of the second position in the planar layout; Based on the planar position, the target indicator is presented in the voice interaction interface.

9. The method according to claim 1, wherein: The adjusting the visual presentation attribute of the user indicator based on the portion of the real-time audio stream corresponding to the user comprises: The size of the user indicator is continuously adjusted based on the accumulated number of words of the text information corresponding to the user in the real-time audio stream, wherein the size is positively correlated with the accumulated number of words.

10. The method according to claim 9, further comprising: In response to the word count accumulation being greater than or equal to an accumulation threshold, continuously adjusting the size of the user indicator is stopped.

11. The method according to claim 1, wherein: The adjusting the visual presentation attribute of the user indicator based on the portion of the real-time audio stream corresponding to the user comprises: The size of the user indicator is adjusted based on the volume of the portion of the real-time audio stream corresponding to the user, wherein the size is positively correlated with the volume.

12. The method according to claim 1, wherein: The adjusting the visual presentation attribute of the user indicator based on the portion of the real-time audio stream corresponding to the user comprises: Based on the accumulated number of words in the text information corresponding to the user in the real-time audio stream, the presentation position of the user indicator is continuously adjusted to move toward the target indicator, wherein the moving distance of the presentation position of the user indicator is positively correlated with the size of the accumulated number of words.

13. The method according to claim 12, further comprising: In response to the presentation position of the user indicator being the same as the presentation position of the target indicator, continuously adjusting the presentation position of the user indicator toward the target indicator is stopped.

14. The method according to claim 12, further comprising: In response to an overlap between the user indicator and the target indicator, continuously adjusting the presentation position of the user indicator to move toward the target indicator is stopped.

15. The method according to claim 1, wherein: The adjusting the visual presentation attribute of the user indicator based on the portion of the real-time audio stream corresponding to the user comprises: In response to determining that the user is currently speaking based on the portion of the real-time audio stream corresponding to the user, the visual style of the user indicator is adjusted to a dynamic icon.

16. The method according to claim 1, further comprising: In response to determining that the user is currently speaking based on a portion of the real-time audio stream corresponding to the user, a dynamic indicator is presented between the user indicator and the target indicator, starting from a presentation position of the user indicator and pointing to a presentation position of the target indicator.

17. The method according to claim 16, wherein: The dynamic indicator includes a dynamic text flow generated based on text information of the content currently being spoken by the user.

18. The method of claim 1, further comprising: In the voice interaction interface, a text box corresponding to the user is added, wherein the text box is used to present text information of a portion of the real-time audio stream corresponding to the user.

19. The method according to any one of claims 1 to 18, wherein: The second position includes: a position where a sound collection device is set in the physical environment, or a position where a target user is located in the physical environment.

20. A speech interaction device based on a large language model, comprising: A user identification and positioning unit, configured to determine a user included in the physical environment and a first position of the user in the physical environment based on a real-time audio stream collected in the physical environment; a user indicator presenting unit configured to present, in a voice interaction interface presented for the physical environment, a user indicator corresponding to the user in association with a target indicator, wherein a relative positional relationship between the user indicator and the target indicator is determined based on a relative positional relationship between the first position and a second position corresponding to the target indicator in the physical environment; The indicator adjustment unit is configured to adjust a visual presentation attribute of the user indicator based on a portion of the real-time audio stream corresponding to the user.

21. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the speech interaction method based on a large language model as described in any one of claims 1-19.

22. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the speech interaction method based on a large language model as described in any one of claims 1-19.

23. A computer program product, comprising a computer program, which, when executed by a processor, implements the speech interaction method based on a large language model according to any one of claims 1-19.