Method, device, electronic device and storage medium for detecting user conversation status
By detecting the conversation time between the main driver and the co-driver user and combining visual information, the problem of false triggering in voice interaction is solved, achieving more accurate voice command filtering and a comfortable user experience, especially in voice control scenarios in the vehicle.
Patent Information
- Application Number
- CN202210214406.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-03
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2042-03-03
AI Technical Summary
During the voice interaction process, due to the limitations of natural language generalization capabilities, the content of conversation between the master and the passenger driver can easily lead to the false triggering of voice commands, resulting in inaccurate voice interaction and poor user experience.
By detecting the conversation time between the first user and the second user, determining whether the two users are in the conversation state based on preset rules, using a multimodal voice endpoint detection model combining voice data and visual information to filter voice commands and adjust music volume.
It effectively avoids the false triggering of voice interaction, improves the accuracy of voice interaction, and provides a comfortable user experience, especially in voice control scenarios in the vehicle.
Smart Images

Figure CN114550721B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to voice interaction technology, and in particular to a method, device, electronic device and storage medium for detecting a user's conversation status. Background Art
[0002] With the rapid development of voice recognition technology, voice interaction has been applied in various fields. For example, in vehicles, the main driver or the front passenger can control in-vehicle applications or devices through voice interaction, which greatly improves the user experience. However, during the voice interaction process, due to the limitations of the generalization ability of natural language, it is easy to cause false triggering of voice commands. For example, if the conversation between the main driver and the front passenger involves relevant voice commands, a response to the voice command will be given, but in fact the response is not what the user needs, resulting in inaccurate voice interaction and a poor user experience. Summary of the Invention
[0003] In order to solve the above-mentioned technical problems such as false triggering caused by the conversation between the driver and the co-driver, the present disclosure is proposed. The embodiments of the present disclosure provide a method, device, electronic device and storage medium for detecting the user conversation state.
[0004] According to one aspect of an embodiment of the present disclosure, a method for detecting a user conversation status is provided, including: determining a first conversation duration of a first user within a first preset time period based on first voice data; determining a second conversation duration of a second user within the first preset time period based on second voice data; determining a total conversation duration within the first preset time period based on the first voice data and the second voice data; determining whether the first conversation duration, the second conversation duration and the total conversation duration meet a preset condition based on a preset rule; and determining that the first user and the second user are in a conversation status in response to the first conversation duration, the second conversation duration and the total conversation duration meeting the preset condition.
[0005] According to another aspect of an embodiment of the present disclosure, a device for detecting a user conversation status is provided, including: a first processing module for determining a first conversation duration of a first user within a first preset time period based on first voice data; a second processing module for determining a second conversation duration of a second user within the first preset time period based on second voice data; a third processing module for determining a total conversation duration within the first preset time period based on the first voice data and the second voice data; a fourth processing module for determining whether the first conversation duration, the second conversation duration and the total conversation duration meet preset conditions based on preset rules; and a fifth processing module for determining that the first user and the second user are in a conversation status in response to the first conversation duration, the second conversation duration and the total conversation duration meeting preset conditions.
[0006] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the storage medium stores a computer program, and the computer program is used to execute the method for detecting the user conversation status described in any of the above embodiments of the present disclosure.
[0007] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing executable instructions of the processor; and the processor for reading the executable instructions from the memory and executing the instructions to implement the method for detecting user conversation status as described in any of the above embodiments of the present disclosure.
[0008] Based on the method, device, electronic device and storage medium for detecting user conversation status provided by the above-mentioned embodiments of the present disclosure, the respective conversation duration and the total conversation duration of the two users are determined based on the voice data of the two users, and corresponding rules are set based on the relevant characteristics of the conversation scene, so that it is possible to determine whether the two users are in a conversation state based on whether the respective conversation duration of the users and the total conversation duration meet the preset conditions, thereby achieving effective detection of the conversation state, which can be applied to filtering voice commands generated in user conversation scenes in voice interaction, avoiding false triggering, and improving the accuracy of voice interaction. It can also be applied to scenarios such as actively lowering the music volume when users are talking to provide users with comfortable music, effectively improving user experience.
[0009] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and examples. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other purposes, features, and advantages of the present disclosure will become more apparent through a more detailed description of the embodiments of the present disclosure in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and are not intended to limit the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.
[0011] Figure 1 This is an exemplary application scenario of the method for detecting user conversation status provided by the present disclosure;
[0012] Figure 2 is a flowchart of a method for detecting a user conversation state provided by an exemplary embodiment of the present disclosure;
[0013] Figure 3 is a flowchart of a method for detecting a user conversation state provided by another exemplary embodiment of the present disclosure;
[0014] Figure 4is a flowchart of step 2011 provided by an exemplary embodiment of the present disclosure;
[0015] Figure 5 is a flowchart of step 2011 provided by another exemplary embodiment of the present disclosure;
[0016] Figure 6 is a flowchart of step 2021 provided by an exemplary embodiment of the present disclosure;
[0017] Figure 7 is a flowchart of a method for detecting a user conversation state provided by yet another exemplary embodiment of the present disclosure;
[0018] Figure 8 is a flowchart of a method for detecting a user conversation state provided by yet another exemplary embodiment of the present disclosure;
[0019] Figure 9 1 is a schematic structural diagram of a device for detecting a user conversation state provided by an exemplary embodiment of the present disclosure;
[0020] Figure 10 is a structural diagram of a device for detecting a user conversation state provided by another exemplary embodiment of the present disclosure;
[0021] Figure 11 is a structural diagram of a first processing unit 511 provided by an exemplary embodiment of the present disclosure;
[0022] Figure 12 is a structural diagram of a second processing unit 521 provided by an exemplary embodiment of the present disclosure;
[0023] Figure 13 is a structural diagram of a first processing module 51 provided by an exemplary embodiment of the present disclosure;
[0024] Figure 14 is a structural diagram of a second processing module 52 provided by an exemplary embodiment of the present disclosure;
[0025] Figure 15 is a structural diagram of a device for detecting a user conversation state provided by yet another exemplary embodiment of the present disclosure;
[0026] Figure 16 is a structural diagram of a fourth processing module 54 provided by an exemplary embodiment of the present disclosure;
[0027] Figure 17 It is a structural diagram of an application embodiment of the electronic device disclosed in the present invention. DETAILED DESCRIPTION
[0028] Below, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0029] It should be noted that the relative arrangement of components and steps, the numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present disclosure unless specifically stated otherwise.
[0030] Those skilled in the art will understand that the terms "first" and "second" in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, and do not represent any specific technical meanings, nor do they indicate a necessary logical order between them.
[0031] It should also be understood that in the embodiments of the present disclosure, “a plurality of” may refer to two or more than two, and “at least one” may refer to one, two, or more than two.
[0032] It should also be understood that any component, data or structure mentioned in the embodiments of the present disclosure can generally be understood as one or more, unless explicitly limited or otherwise indicated in the context.
[0033] In addition, the term "and / or" in this disclosure is merely a description of the association relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this disclosure generally indicates that the related objects are in an "or" relationship.
[0034] It should also be understood that the description of the various embodiments in this disclosure focuses on the differences between the various embodiments, and the same or similar aspects thereof can be referenced with each other. For the sake of brevity, they will not be described one by one.
[0035] At the same time, it should be understood that for the convenience of description, the sizes of the various parts shown in the drawings are not drawn according to the actual proportional relationship.
[0036] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.
[0037] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0038] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0039] The embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate in conjunction with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, among others.
[0040] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, where tasks are performed by remote processing devices linked via a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media, including storage devices.
[0041] Overview of the Disclosure
[0042] In the process of implementing the present disclosure, the inventors discovered that during the voice interaction process, due to the limitations of the generalization ability of natural language, voice commands are easily triggered incorrectly. For example, if the conversation between the driver and the co-driver involves relevant voice commands, a response to the voice command will be given, but in fact the response is not what the user needs, which makes the voice interaction inaccurate and leads to a poor user experience.
[0043] Exemplary Overview
[0044] Figure 1This is an exemplary application scenario for the user conversation state detection method provided by the present disclosure. Within a vehicle, the driver or passenger can control in-vehicle applications or devices, such as the air conditioning, music, navigation, windows, and doors, through voice interaction. After the corresponding voice acquisition device captures the user's voice, it transmits it to the vehicle's control device, such as an onboard computing platform. The vehicle's control device recognizes the user's voice, determines a voice command based on the user's voice, and sends it to the corresponding execution device of the interaction object to perform the corresponding operation based on the user's voice command. However, the determined voice command may be the content of a conversation between the driver and passenger, rather than a valid voice command. Using the user conversation state detection method disclosed in the present disclosure, it is possible to effectively detect whether the driver and passenger are in a conversation state. This allows filtering of received voice commands, eliminating voice commands generated during a conversation, and preventing false triggering of voice interaction, thereby improving the accuracy of voice interaction. Furthermore, when the driver and passenger are detected to be in a conversation state, the system can proactively reduce the music volume to provide comfortable music without disrupting the conversation, thereby enhancing the user experience.
[0045] The application of the detection of the conversation status of the driver and the co-driver is not limited to the above-mentioned voice command filtering and active reduction of music volume, and can be applied in any other possible aspects, which is not limited in this disclosure.
[0046] The method disclosed herein can also be combined with visual information to detect the user's conversation status to improve the accuracy of the detection results. For example, a camera that monitors the main driver and the front passenger is set on the vehicle to collect image data of the main driver and the front passenger to assist in the detection of the conversation status.
[0047] The disclosed method for detecting user conversation status is not limited to the aforementioned application scenarios. Any scenario requiring voice interaction can use the disclosed method to detect whether the user issuing a voice command is currently conversing with another user. This can include game rooms and other scenarios with devices controllable via voice interaction, though the disclosed embodiments are not limiting.
[0048] Exemplary Methods
[0049] Figure 2 This is a flow chart of a method for detecting a user's conversation status provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to electronic devices, such as servers or terminals, and specifically on vehicle-mounted computing platforms, such as Figure 2 As shown, the following steps are included:
[0050] Step 201: Determine a first conversation duration of a first user within a first preset time period based on first voice data.
[0051] Among them, the first voice data can be the voice data of the first user collected or obtained after collection and processing, or it can include the voice data of the first user and the voice data of other users, and there is no specific limitation. The first preset time period can be any time period in the first voice data and can be set according to actual needs. For example, when used for filtering voice instructions in voice interaction, the first preset time period can be a time period including the voice period corresponding to the first voice instruction. The first conversation duration is the length of time that the first user is in a speaking state within the first preset time period. Since one party speaks and the other party listens when the user is talking, there will be non-voice segments and voice segments in the first voice data. The first conversation duration is the length of the first user's voice segment. The specific determination of the first conversation duration can be carried out in any feasible manner, such as by voice endpoint detection to determine the first conversation duration, or by combining voice data and corresponding visual information to determine the first conversation duration, and this disclosure does not limit it.
[0052] For example, the first voice data is data obtained that only contains the voice of the first user. Voice endpoint detection can be used to detect the voice segments in the first voice data, and the sum of the duration of each voice segment is calculated to be the first conversation duration of the first user.
[0053] For another example, the first voice data includes the voice of the first user and the voices of other users. The voice data of the first user can be extracted from the first voice data in an practicable manner, and then the conversation duration of the first user can be determined.
[0054] Step 202: Determine a second conversation duration of the second user within a first preset time period based on the second voice data.
[0055] The second conversation duration is determined similarly to the first conversation duration and will not be further described here. Similar to the first voice data, the second voice data may be collected or processed voice data of the second user, or may include the second user's voice and the voices of other users, without limitation.
[0056] For example, in a vehicle, to determine whether the main driver (first user) and the co-driver (second user) are in a conversation state, the main driver's voice acquisition device can be used to collect the main driver's first voice data, and the co-driver's voice acquisition device can be used to collect the co-driver's second voice data. However, since the space inside the vehicle is relatively small, the main driver's voice acquisition device can collect the co-driver's voice at the same time, and the co-driver's voice acquisition device can also collect the main driver's voice at the same time. In this case, the collected voice data can be used to first determine the first voice data that mainly includes the main driver's voice and the second voice data that mainly includes the co-driver's voice. For example, the voice volume of the voice data collected by the two voice acquisition devices in the same time period can be used to determine whether it is the main driver's voice or the co-driver's voice. Since the main driver's voice acquisition device is close to the main driver and the co-driver's voice acquisition device is far away from the main driver, the main driver's voice collected by the two devices in the same time period has different volume. The specific principles will not be repeated here. Regardless of the method used, as long as the first voice data of the first user and the second voice data of the second user can be determined, it will be sufficient.
[0057] The specific voice collection method and the number of voice collection devices can be set according to actual needs, and are not limited to the above-mentioned method of two voice collection devices collecting the first voice data of the main driver user and the second voice data of the front passenger user respectively.
[0058] For example, a voice collection device may be used to simultaneously collect the first voice data of the first user and the second voice data of the second user, and the first voice data may be separated from the second voice data by a certain distinguishing method, for example, separation by voiceprint recognition. The specific separation method may be any feasible method, which is not limited in this disclosure.
[0059] In order to determine the conversation duration of the first user and the second user respectively, time information can be carried in the first voice data and the second voice data to determine the duration of the voice segments therein, and then determine the first conversation duration of the first user and the second conversation duration of the second user.
[0060] Step 203: Determine the total conversation duration within the first preset time period based on the first voice data and the second voice data.
[0061] Among them, the total conversation duration refers to the total duration that the first user and the second user are in a voice activation state (i.e., a speaking state) within a first preset time period. The union of the first voice data and the second voice data can be taken, and based on voice endpoint detection, the total duration of the speaking state within the first preset time period can be determined and combined as the total conversation duration within the first preset time period.
[0062] In actual applications, when determining the total conversation duration based on the first and second voice data, the total conversation duration can be determined by taking the union of the time periods corresponding to the voice segments obtained by endpoint detection of the first and second voice data based on the intermediate results of processing the first and second voice data in the aforementioned steps, such as the endpoint detection results of the first and second voice data. This can be set based on actual needs.
[0063] Steps 201-203 are in no particular order.
[0064] Step 204: Determine based on preset rules whether the first conversation duration, the second conversation duration, and the total conversation duration meet preset conditions.
[0065] Preset rules and conditions can be set based on the characteristics of user conversations. For example, during a user conversation, users may take turns speaking. Therefore, within a first preset time period, the first and second users alternate speaking. Consequently, the total conversation duration may account for a larger proportion of the first preset time period, with the first user's non-speech periods being supplemented by the second user's speech periods. For another example, within the total conversation duration, the first user's first conversation duration and the second user's second conversation duration may each account for a certain proportion. Based on these characteristics, preset rules can be set to perform certain calculations, such as calculating a proportion, and preset conditions can be set, such as whether the proportion is greater than a threshold. Based on the preset rules, it can then be determined whether the first conversation duration, the second conversation duration, and the total conversation duration meet the preset conditions. The specific configuration can be tailored to actual needs and will not be elaborated upon here.
[0066] Step 205: In response to the first conversation duration, the second conversation duration, and the total conversation duration satisfying a preset condition, it is determined that the first user and the second user are in a conversation state.
[0067] When the first conversation duration, the second conversation duration, and the total conversation duration meet a preset condition, it indicates that the first user and the second user are in a conversation state.
[0068] The method for detecting user conversation status provided in this embodiment can determine the respective conversation durations and total conversation durations of two users based on their voice data. Based on the characteristics of the conversation scenario, corresponding rules are set to determine whether two users are in a conversation state based on whether the respective conversation durations and the total conversation duration meet preset conditions. This effectively detects the conversation state and can be applied to filter voice commands generated in user conversation scenarios during voice interaction, preventing false triggers and improving the accuracy of voice interaction. It can also be used to proactively lower the music volume during a conversation to provide users with comfortable music without disrupting the conversation, effectively improving the user experience.
[0069] Figure 3It is a flowchart of a method for detecting a user conversation state provided by another exemplary embodiment of the present disclosure.
[0070] In an optional example, step 201 may specifically include:
[0071] Step 2011: Determine a first conversation duration of a first user based on the first voice data and the corresponding first image data.
[0072] Among them, the first image data can be image data of the first user collected based on an image acquisition device (such as a camera), and can be an image or a video, with no specific limitation. For example, a camera for monitoring the main driver and the co-driver is provided on the vehicle, and the first image data corresponding to the main driver user and the second image data corresponding to the co-driver user can be obtained based on the positional relationship between the shooting area and the main driver area and the co-driver area. The first conversation duration is determined in combination with the first voice data and the first image data, which can be specifically implemented based on a multimodal voice endpoint detection model that combines voice and vision. The multimodal voice endpoint detection model is a model that combines voice data and visual information for voice endpoint detection. Since the information of the image frame in the visual information will not undergo additional changes due to the presence of noise, it can help to exclude paragraphs corresponding to noise content from the voice. Therefore, by referring to the feature information of the voice and the feature information of the vision at the same time, the detection accuracy of the voice activation state can be significantly improved.
[0073] Exemplarily, an audio feature sequence can be obtained based on the first voice data, an image feature sequence can be obtained based on the first image data, and a fused feature sequence can be obtained based on the audio feature sequence and the image feature sequence. The fused feature sequence can be input into a multimodal voice endpoint detection model to obtain a predicted probability sequence, and the first detection result described above can be obtained based on the predicted probability sequence. The predicted probability sequence includes the probability values of the audio content at each time point in the first voice data belonging to each type. According to the size of the predicted probability value, by setting a probability threshold, it can be determined whether the audio content at the corresponding time point is in a voice activation state. The specific principles will not be repeated here.
[0074] The multimodal speech endpoint detection model is essentially a classification model and can be implemented using any feasible neural network model, which is not limited in this disclosure.
[0075] In an optional example, step 202 may specifically include:
[0076] Step 2021: Determine a second conversation duration of the second user based on the second voice data and the corresponding second image data.
[0077] The principle for determining the second conversation duration by combining voice and image is similar to that for the first conversation duration and will not be described in detail here.
[0078] In an optional example, step 203 may specifically include:
[0079] Step 2031 : Determine the total conversation duration within a first preset time period based on the first voice data, the first image data corresponding to the first voice data, the second voice data, and the second image data corresponding to the second voice data.
[0080] Specifically, the time period of the first user's voice activation state can be determined based on the first voice data and the first image data, and the time period of the second user's voice activation state can be determined based on the second voice data and the second image data. The total conversation duration can be determined by combining the union of the two voice activation state time periods.
[0081] For example, if the first preset time period is 0-10 seconds, the first user's voice active time periods are 1-3 seconds and 6-8 seconds, and the second user's voice active time periods are 4-5 seconds and 9-10 seconds, then the first conversation duration is 4 seconds, the second conversation duration is 2 seconds, and the total conversation duration is 6 seconds. In some cases, the first and second users may speak simultaneously, causing their voice active time periods to overlap. In other cases, there may be voice segments from other users within the first preset time period, resulting in the sum of the first and second conversation durations being less than the total conversation duration.
[0082] The present disclosure can effectively improve the accuracy of the conversation duration by combining voice and image to determine the first conversation duration, the second conversation duration and the total conversation duration.
[0083] In an alternative example, Figure 4 2 is a flow chart of step 2011 provided in an exemplary embodiment of the present disclosure. In this example, step 2011 determines the first conversation duration of the first user based on the first voice data and the corresponding first image data, including:
[0084] Step 20111: Take each speech frame in the first speech data as the current speech frame, take the image frame corresponding to the current speech frame in the first image data as the current image frame, and determine the first speech activation state corresponding to the current frame based on the current speech frame and the current image frame.
[0085] An image frame is a temporal unit of the collected image data, with each image frame corresponding to a point in time. A voice frame is a temporal unit of the collected voice data, with each voice frame also corresponding to a point in time. The first voice activation state can include active and inactive states. Active indicates that the user is speaking, while inactive indicates that the user is not speaking.
[0086] The first image data includes the action image of the user who makes the speech, so the image feature information obtained by feature extraction based on the image frame can be used to represent whether the user is speaking at a certain point in time. For example, when the extracted image feature information shows that the user has special facial movements (such as the mouth is open), it can be considered that the user is in a speaking state at that time point; and when the extracted image feature information shows that the user does not have special facial movements, it can be considered that the user is not in a speaking state at that time point.
[0087] The first voice data includes voice content and noise content uttered by the user. For a voice segment to be detected, the voice content and the noise content are mixed together; while a non-voice segment only includes noise content.
[0088] The first image data and the first voice data are combined to detect the voice activity state and obtain the first voice activity state corresponding to the current frame. Even if the voice data is collected in a complex scene with high noise, the information content of the image frame is not affected by the noise. The image frame information includes the user's movements related to the speaking state at the image level. Therefore, by referencing the features of the image frame, non-voice segments caused by noise can be eliminated, thereby improving the accuracy of voice activity state detection.
[0089] Optionally, in order to further improve the accuracy, the image feature sequence and audio feature sequence corresponding to the current frame can also be obtained, and the voice activation state of the current frame can be determined by combining the features of the current frame with the historical features. Specifically, the image feature sequence can be obtained by extracting the features of the current image frame and the historical image frames of a preset number of frames before the current image frame. The image feature sequence includes multiple image feature information at multiple time points that are continuous on the time scale, and the multiple image feature information are respectively extracted from multiple image frames corresponding to the multiple time points. The image feature sequence represents the action images related to the speaking state presented by the user at the image level at multiple time points. The audio feature sequence includes multiple audio feature information at multiple time points that are continuous on the time scale, and the multiple audio feature information are respectively extracted from multiple voice frames corresponding to the multiple time points. The audio feature sequence represents the audio features of the voice data at multiple time points. The specific method of extracting the image feature sequence and the audio feature sequence is not limited in this disclosure.
[0090] Step 20112, determine whether the current cycle is finished.
[0091] Among them, the current cycle refers to the cycle of the time point corresponding to the current frame, and the cycle can be set according to actual needs, such as starting from the starting frame, and every 0.5 seconds as a cycle. This example can be to process the collected data in real time during the collection of the first image data and the first voice data, or it can be to process it frame by frame after the collection is completed, and there is no specific limitation. For example, when the first voice data and the first image data have been obtained, each voice frame can be traversed from the starting frame, and each voice frame can be used as the current voice frame. The current image frame is determined based on the time information. After the first voice activation state corresponding to the current frame is determined based on the current voice frame and the current image frame, it is determined whether the current cycle is over based on the correspondence between the frame and the duration. For example, each frame is 10 milliseconds, and whether the current cycle is over can be determined based on the number of frames processed in the current cycle and the cycle duration. The specific determination method will not be repeated.
[0092] Step 20113: In response to the end of the current cycle, determine the first activation duration in the current cycle based on the first voice activation state corresponding to each frame in the current cycle, and enter the processing of the next cycle until the processing of the first preset time period is completed.
[0093] After the current cycle ends, the first activation duration of the first user in the current cycle can be determined based on the first voice activation state corresponding to each frame in the current cycle. The first activation duration of the current cycle can be stored and processed in the next cycle. Similarly, the first activation duration of the first user in each cycle in the first preset time period can be obtained.
[0094] Exemplarily, the first preset time period is set to a 30-second time period, that is, the first voice data is 30 seconds long, or the 30-second content in the first voice data is processed, which can be set according to actual needs. For example, processing is performed frame by frame, and each 0.5 second is a cycle, then there are a total of 60 cycles, and each cycle records a first activation duration, such as 0.1 seconds for the first cycle, 0.3 seconds for the second cycle, and so on. It can be set according to actual needs. Taking a cycle as an example, for example, there are a total of 100 frames, each frame has a corresponding first voice activation state, and the first activation duration in the cycle can be obtained by counting the time that the frames in which the first voice activation state is activated in the cycle are activated.
[0095] Step 20114: Determine a first conversation duration based on the first activation duration in each cycle within the first preset time period.
[0096] Specifically, the first conversation duration is the sum of the first activation durations of each cycle.
[0097] Step 20115: In response to the current cycle not ending, record the first voice activation state corresponding to the current frame and proceed to the processing of the next frame until the current cycle ends.
[0098] If the current cycle has not ended, the first voice activation state corresponding to the current frame is recorded. The specific recording method can be set according to actual needs, such as storing in a memory, or recording through a FIFO queue, which is not specifically limited.
[0099] The present disclosure determines the user's conversation duration by processing frame by frame, and can realize real-time processing of voice data and image data, thereby improving the real-time performance of detection.
[0100] In an optional example, step 20113 specifically includes: in response to the end of the current cycle, determining the first activation duration in the current cycle based on the first voice activation state corresponding to each frame in the current cycle, adding the first activation duration in the current cycle to the first queue, and entering the processing of the next cycle until the processing of the first preset time period is completed.
[0101] The first queue may be a first-in-first-out (FIFO) queue. After determining the first activation duration in the current cycle, the first activation duration is added to the first queue for recording to facilitate subsequent determination of the first conversation duration.
[0102] Accordingly, the first conversation duration determined based on the first activation duration in each cycle within the first preset time period in step 20114 includes: determining the first conversation duration based on the first activation duration in each cycle in the first queue.
[0103] The specific principles can be found in the above content and will not be repeated here.
[0104] The present disclosure records the activation duration of each cycle through a queue, which has low overhead and simple processing.
[0105] In an alternative example, Figure 5 2 is a flow chart of step 2011 provided by another exemplary embodiment of the present disclosure. In this example, after step 20111, the following steps are further included:
[0106] Step 20111a: In response to the first speech activation state corresponding to the current frame being active, the lip movement state corresponding to the current image frame is determined based on the trained lip movement detection model.
[0107] The lip movement detection model can employ any feasible neural network model. Lip movement states can include dynamic and static states, where dynamic indicates lip movement and static indicates no lip movement. The specific detection principle, for example, can be to determine the lip movement state by combining the lip features of the current image frame with those of at least one historical image frame. For example, if the user's lips change in the image frame sequence including the current frame, the user's lip movement state in the current frame is determined to be dynamic. The specific principle can be configured based on actual needs and is not limited to the above principle, and will not be elaborated here.
[0108] Step 20111b: In response to the lip movement state corresponding to the current image frame being dynamic, determining that the first voice activation state corresponding to the current frame is active.
[0109] When it is determined that the lip movement state of the first user in the current image frame is dynamic, it means that the first user is indeed speaking, so it can be determined that the first voice activation state of the current frame is indeed activated.
[0110] Step 20111c: In response to the lip movement state corresponding to the current image frame being static, determining that the first voice activation state corresponding to the current frame is inactive.
[0111] It is currently determined that the lip movement state of the first user in the current image frame is static, indicating that the first user is not speaking. Since the current image frame is only an auxiliary voice frame for detection when the first voice activation state of the current frame is determined based on the current voice frame and the current image frame, this example further determines the first voice activation state in combination with pure visual lip movement detection to further improve the accuracy of the detection results.
[0112] In an alternative example, Figure 6 2 is a flow chart of step 2021 provided in an exemplary embodiment of the present disclosure. In this example, step 2021 determines the second conversation duration of the second user based on the second voice data and the corresponding second image data, including:
[0113] Step 20211: Take each voice frame in the second voice data as the current voice frame, take the image frame corresponding to the current voice frame in the second image data as the current image frame, and determine the second voice activation state corresponding to the current frame based on the current voice frame and the current image frame.
[0114] Step 20212, determine whether the current cycle has ended.
[0115] Step 20213, in response to the end of the current cycle, determine the second activation duration in the current cycle based on the second voice activation state corresponding to each frame in the current cycle, and enter the processing of the next cycle until the processing of the first preset time period is completed.
[0116] Step 20214: Determine a second conversation duration based on the second activation durations in each cycle within the first preset time period.
[0117] Step 20215: In response to the current cycle not ending, record the second voice activation state corresponding to the current frame and proceed to the processing of the next frame until the current cycle ends.
[0118] The specific operations of the above steps 20211-20215 are similar to the above steps 20111-20115 and will not be repeated here.
[0119] In an optional example, step 20213 determines the second activation duration in the current cycle based on the second voice activation state corresponding to each frame in the current cycle in response to the end of the current cycle, including: determining the second activation duration in the current cycle based on the second voice activation state corresponding to each frame in the current cycle in response to the end of the current cycle, adding the second activation duration in the current cycle to the second queue, and entering the processing of the next cycle until the processing of the first preset time period is completed.
[0120] The specific operation of this step is shown in the above content and will not be repeated here.
[0121] Accordingly, the second conversation duration determined based on the second activation duration in each cycle within the first preset time period in step 20214 includes: determining the second conversation duration based on the second activation duration in each cycle in the second queue.
[0122] Figure 7 It is a flowchart of a method for detecting a user conversation state provided by yet another exemplary embodiment of the present disclosure.
[0123] In an optional example, determining the first conversation duration of the first user based on the first voice data and the corresponding first image data in step 2011 includes:
[0124] Step 2011a: Determine a first fusion feature based on the first speech data and the first image data.
[0125] The first fusion feature may include fusion features corresponding to each frame of the first voice data and the first image data, respectively. The fusion feature corresponding to a frame may be a fusion result of the image feature and audio feature of the frame. To improve accuracy, it may also be a fusion result of the image feature sequence and audio feature sequence obtained from the frame and a preset number of historical frames preceding the frame. The specific setting can be based on actual needs. The extraction of image features and audio features may be implemented in any feasible manner and is not limited by this disclosure.
[0126] Step 2011b: Detect the first fusion feature based on the pre-trained multimodal speech endpoint detection model to obtain a first detection result. The first detection result includes first probability values corresponding to multiple time points within a first preset time period.
[0127] Among them, the multimodal speech endpoint detection model is a model that combines speech data and visual information to perform speech endpoint detection. Since the information of the image frame in the visual information will not undergo additional changes due to the presence of noise, it can help to exclude paragraphs corresponding to noise content from the speech. Therefore, by referring to the feature information of speech and the feature information of vision at the same time, the detection accuracy of the speech activation state can be significantly improved. The multimodal speech endpoint detection model can adopt any feasible classification model, which is not limited by this disclosure. The first probability value represents the probability that the speech frame at the corresponding time point is in the speech activation state.
[0128] Step 2011c: Determine a first conversation duration of the first user based on the first detection result.
[0129] Because the first detection result includes first probability values corresponding to multiple time points within the first preset time period, and the first probability values represent the probability that the speech frame at the corresponding time point is in the voice activation state, the type of voice activation state corresponding to each time point can be determined by setting a probability threshold and comparing the first probability values with the probability threshold. Based on the type of voice activation state corresponding to each time point, the time point of the active type can be determined, and the first conversation duration can be calculated based on the active type of the time point.
[0130] In an optional example, determining the second conversation duration of the second user based on the second voice data and the corresponding second image data in step 2021 includes:
[0131] Step 2021a: Determine a second fusion feature based on the second speech data and the second image data.
[0132] Step 2021b: Detect the second fusion feature based on the pre-trained multimodal speech endpoint detection model to obtain a second detection result. The second detection result includes second probability values corresponding to multiple time points within the first preset time period.
[0133] Step 2021c: Determine a second conversation duration of the second user based on the second detection result.
[0134] The specific operations of steps 2021a-2021c are similar to those of steps 2011a-2011c and will not be repeated here.
[0135] In an optional example, determining the total conversation duration based on the first voice data and the corresponding first image data, the second voice data and the corresponding second image data in step 2031 may specifically include:
[0136] 20311. Determine the total conversation duration based on the first detection result and the second detection result.
[0137] Exemplarily, the union of the time periods corresponding to the first detection result and the second detection result can be taken, and the total conversation duration can be determined based on the union. The total conversation duration refers to the time period within the first preset time period, whether it is the first user or the second user or even other users, as long as it is a voice segment, it will be included in the total conversation duration.
[0138] The present disclosure detects the user's voice activation status through a multimodal voice endpoint detection model, and then determines the user's conversation duration based on the voice activation status at each time point, effectively improving the accuracy of the detection results.
[0139] In an optional example, step 2011c determines the first conversation duration of the first user based on the first detection result, including: determining the first time point in the conversation state among each time point based on the first probability values corresponding to multiple time points within the first preset time period; determining the first conversation duration based on the first time point in the conversation state among each time point.
[0140] Among them, the type of the conversation state, that is, the voice activation state is activation, and a probability threshold can be set. The time point when the first probability value is greater than the probability threshold is taken as the first time point in the conversation state. The specific probability threshold can be set according to actual needs, and this disclosure does not limit it.
[0141] In an optional example, step 2021c determines the second conversation duration of the second user based on the second detection result, including: determining the second time point in the conversation state among each time point based on the second probability values corresponding to multiple time points within the first preset time period; determining the second conversation duration based on the second time point in the conversation state among each time point.
[0142] The specific operations of this step are similar to the operations related to the first conversation duration. Please refer to the above content for details and will not be repeated here.
[0143] Figure 8 It is a flowchart of a method for detecting a user conversation state provided by yet another exemplary embodiment of the present disclosure.
[0144] In an optional example, before determining the first conversation duration of the first user within the first preset time period based on the first voice data in step 201, the method of the present disclosure further includes:
[0145] Step 301: Determine a first state of a first user based on first image data corresponding to first voice data.
[0146] The first state may include a calling state and a non-calling state. Determining the first state of the first user based on the first image data can be achieved in any feasible manner, such as through image classification. Specifically, for example, a pre-trained neural network classification model for detecting calling states is used to classify the first image data to determine the first state of the first user. The specific network structure of the model can adopt any feasible structure and is not limited by this disclosure.
[0147] Step 302 : In response to the first state being a non-calling state, determining a second state of the second user based on second image data corresponding to second voice data.
[0148] The principle of determining the second state is similar to that of the first user. Please refer to the above content for details and will not be repeated here.
[0149] Accordingly, step 201 of determining a first conversation duration of the first user within a first preset time period based on the first voice data includes:
[0150] In step 201a, in response to the second state being a non-calling state, a first conversation duration of the first user within a first preset time period is determined based on the first voice data.
[0151] Specifically, when the first user or the second user is in a call state, it means that the first user and the second user are not talking. Therefore, it can be determined that the first user and the second user are not in a conversation state, and there is no need to perform a subsequent conversation duration determination process to reduce the waste of processing resources and improve data processing efficiency.
[0152] It is understandable that the second conversation duration and the total conversation duration are also executed after determining that the second state is a non-calling state, which will not be described in detail here.
[0153] In an optional example, it is also possible to determine whether the first user or the second user is in a calling state, which can be specifically set according to actual needs.
[0154] In an optional example, determining whether the first conversation duration, the second conversation duration, and the total conversation duration meet preset conditions based on preset rules in step 204 includes:
[0155] Step 2041: Determine a first ratio based on the total conversation duration and the total duration of the first preset time period.
[0156] Among them, the first ratio refers to the ratio of total conversation time to total time.
[0157] Step 2042: Determine a second ratio based on the first conversation duration and the total conversation duration.
[0158] The second ratio refers to the ratio of the first conversation duration to the total conversation duration.
[0159] Step 2043: Determine a third ratio based on the second conversation duration and the total conversation duration.
[0160] Among them, the third ratio is the ratio of the second conversation duration to the total conversation duration.
[0161] Step 2044: Based on the first ratio, the second ratio, and the third ratio, determine whether the first conversation duration, the second conversation duration, and the total conversation duration meet a preset condition.
[0162] Specifically, a first threshold, a second threshold, and a third threshold can be set. When the first ratio is greater than the first threshold, the second ratio is greater than the second threshold, and the third ratio is greater than the third threshold, it can be determined whether the first conversation duration, the second conversation duration, and the total conversation duration meet the preset conditions. The specific first threshold, second threshold, and third threshold can be set according to actual needs, for example, the first threshold is 80%, the second threshold is 10%, and the third threshold is 10%. Other values can also be set, which will not be detailed here.
[0163] The above-mentioned embodiments and optional examples of the present disclosure may be implemented separately or in any combination without conflict, and the present disclosure does not limit them.
[0164] Any of the user conversation status detection methods provided in the embodiments of the present disclosure can be executed by any appropriate device with data processing capabilities, including but not limited to terminal devices and servers. Alternatively, any of the user conversation status detection methods provided in the embodiments of the present disclosure can be executed by a processor, such as by invoking corresponding instructions stored in a memory to execute any of the user conversation status detection methods mentioned in the embodiments of the present disclosure. This will not be further described below.
[0165] Exemplary devices
[0166] Figure 9 FIG is a schematic diagram of a device for detecting a user conversation state provided by an exemplary embodiment of the present disclosure. The device of this embodiment can be used to implement the corresponding method embodiments of the present disclosure, such as Figure 9 The device shown includes a first processing module 51 , a second processing module 52 , a third processing module 53 , a fourth processing module 54 and a fifth processing module 55 .
[0167] The first processing module 51 is used to determine the first conversation duration of the first user within the first preset time period based on the first voice data; the second processing module 52 is used to determine the second conversation duration of the second user within the first preset time period based on the second voice data; the third processing module 53 is used to determine the total conversation duration within the first preset time period based on the first voice data and the second voice data; the fourth processing module 54 is used to determine whether the first conversation duration, the second conversation duration and the total conversation duration meet the preset conditions based on preset rules; the fifth processing module 55 is used to determine that the first user and the second user are in a conversation state in response to the first conversation duration, the second conversation duration and the total conversation duration meeting the preset conditions.
[0168] Figure 10 It is a structural diagram of a device for detecting a user conversation state provided by another exemplary embodiment of the present disclosure.
[0169] In an optional example, the first processing module 51 includes: a first processing unit 511, configured to determine a first conversation duration of the first user based on the first voice data and the corresponding first image data.
[0170] In an optional example, the second processing module 52 includes: a second processing unit 521, configured to determine a second conversation duration of the second user based on the second voice data and the corresponding second image data.
[0171] In an optional example, the third processing module 53 includes: a third processing unit 531, which is used to determine the total conversation time within the first preset time period based on the first voice data, the first image data corresponding to the first voice data, the second voice data, and the second image data corresponding to the second voice data.
[0172] In an alternative example, Figure 11 FIG2 is a schematic diagram of the structure of a first processing unit 511 provided by an exemplary embodiment of the present disclosure. In this example, the first processing unit 511 includes: a first determining subunit 5111, a second determining subunit 5112, a first processing subunit 5113, a third determining subunit 5114, and a second processing subunit 5115.
[0173] The first determining subunit 5111 is configured to use each speech frame in the first speech data as a current speech frame and the image frame in the first image data corresponding to the current speech frame as a current image frame, and determine a first speech activation state corresponding to the current frame based on the current speech frame and the current image frame. The second determining subunit 5112 is configured to determine whether a current cycle has ended. The first processing subunit 5113 is configured to, in response to the second determining subunit 5112 determining that the current cycle has ended, determine a first activation duration within the current cycle based on the first speech activation state corresponding to each frame within the current cycle, and proceed to processing the next cycle until processing of the first preset time period is completed. The third determining subunit 5114 is configured to determine a first conversation duration based on the first activation duration of each cycle within the first preset time period obtained by the first processing subunit 5113. The second processing subunit 5115 is configured to, in response to the second determining subunit 5112 determining that the current cycle has not ended, record the first speech activation state corresponding to the current frame and proceed to processing the next frame until processing of the current cycle ends.
[0174] In an optional example, the first processing unit 511 further includes: a third processing subunit 5116, configured to add the first activation duration in the current cycle to the first queue.
[0175] Accordingly, the third determining subunit 5114 is specifically configured to determine the first conversation duration based on the first activation duration in each cycle in the first queue.
[0176] In an optional example, the first processing unit 511 further includes: a fourth determining subunit 5117 , a fifth determining subunit 5118 , and a sixth determining subunit 5119 .
[0177] The fourth determination subunit 5117 is used to determine the lip movement state corresponding to the current image frame based on the trained lip movement detection model in response to the first voice activation state corresponding to the current frame being activated; the fifth determination subunit 5118 is used to determine that the first voice activation state corresponding to the current frame is activated in response to the lip movement state corresponding to the current image frame being dynamic; the sixth determination subunit 5119 is used to determine that the first voice activation state corresponding to the current frame is inactivated in response to the lip movement state corresponding to the current image frame being static.
[0178] In an alternative example, Figure 12 FIG2 is a schematic diagram of the structure of a second processing unit 521 provided by an exemplary embodiment of the present disclosure. In this example, the second processing unit 521 includes: a fourth processing subunit 5211, a seventh determining subunit 5212, a fifth processing subunit 5213, an eighth determining subunit 5214, and a sixth processing subunit 5215.
[0179] The fourth processing subunit 5211 is configured to use each voice frame in the second voice data as a current voice frame and the image frame in the second image data corresponding to the current voice frame as a current image frame, and determine, based on the current voice frame and the current image frame, the second voice activation state corresponding to the current frame. The seventh determining subunit 5212 is configured to determine whether the current cycle has ended. The fifth processing subunit 5213 is configured to, in response to the fourth processing subunit 5211 determining that the current cycle has ended, determine a second activation duration within the current cycle based on the second voice activation state corresponding to each frame within the current cycle, and proceed to processing the next cycle until processing of the first preset time period is completed. The eighth determining subunit 5214 is configured to determine a second conversation duration based on the second activation duration of each cycle within the first preset time period obtained by the fifth processing subunit 5213. The sixth processing subunit 5215 is configured to, in response to the fourth processing subunit 5211 determining that the current cycle has not ended, record the second voice activation state corresponding to the current frame and proceed to processing the next frame until processing of the current cycle ends.
[0180] In an optional example, the second processing unit 521 also includes: a seventh processing sub-unit 5216, used to add the second activation duration in the current cycle to the second queue; accordingly, the eighth determination sub-unit 5214 is specifically used to: determine the second conversation duration based on the second activation duration in each cycle in the second queue.
[0181] In an optional example, the second processing unit 521 further includes: a ninth determining subunit 5217 , an eighth processing subunit 5218 , and a ninth processing subunit 5219 .
[0182] The ninth determination subunit 5217 is used to determine the lip movement state corresponding to the current image frame based on the trained lip movement detection model in response to the second voice activation state corresponding to the current frame being activated; the eighth processing subunit 5218 is used to determine that the second voice activation state corresponding to the current frame is activated in response to the lip movement state corresponding to the current image frame being dynamic; the ninth processing subunit 5219 is used to determine that the second voice activation state corresponding to the current frame is inactivated in response to the lip movement state corresponding to the current image frame being static.
[0183] In an alternative example, Figure 13 FIG. 5 is a schematic diagram of the structure of a first processing module 51 provided by an exemplary embodiment of the present disclosure. In this example, the first processing module 51 includes: a first determining unit 511a, a first detecting unit 511b, and a second determining unit 511c.
[0184] The first determination unit 511a is used to determine the first fusion feature based on the first voice data and the first image data; the first detection unit 511b is used to detect the first fusion feature obtained by the first determination unit 511a based on a pre-trained multimodal voice endpoint detection model to obtain a first detection result, and the first detection result includes first probability values corresponding to multiple time points within a first preset time period; the second determination unit 511c is used to determine the first conversation duration of the first user based on the first detection result obtained by the first detection unit 511b.
[0185] In an alternative example, Figure 14 FIG2 is a schematic diagram of the structure of a second processing module 52 provided by an exemplary embodiment of the present disclosure. In this example, the second processing module 52 includes: a third determining unit 521a, a second detecting unit 521b, and a fourth determining unit 521c.
[0186] The third determination unit 521a is used to determine the second fusion feature based on the second voice data and the second image data; the second detection unit 521b is used to detect the second fusion feature obtained by the third determination unit 521a based on the pre-trained multimodal voice endpoint detection model to obtain a second detection result, and the second detection result includes second probability values corresponding to multiple time points within the first preset time period; the fourth determination unit 521c is used to determine the second conversation duration of the second user based on the second detection result obtained by the second detection unit 521b.
[0187] Similarly, the third processing module 53 can also perform the unit division as the first processing module 51 to obtain the total conversation duration, which will not be repeated here.
[0188] In an optional example, the second determination unit 511c is specifically used to: determine the first time point in the conversation state among each time point based on the first probability values corresponding to multiple time points within the first preset time period; and determine the first conversation duration based on the first time point in the conversation state among each time point.
[0189] In an optional example, the fourth determination unit 521c is specifically used to: determine the second time point in the conversation state among each time point based on the second probability values corresponding to multiple time points within the first preset time period; determine the second conversation duration based on the second time point in the conversation state among each time point.
[0190] Figure 15 It is a structural diagram of a device for detecting a user conversation state provided by yet another exemplary embodiment of the present disclosure.
[0191] In an optional example, the apparatus of the present disclosure further includes: a first determination module 56 and a second determination module 57 .
[0192] The first determination module 56 is used to determine the first state of the first user based on the first image data corresponding to the first voice data; the second determination module 57 is used to determine the second state of the second user based on the second image data corresponding to the second voice data in response to the first state being a non-calling state; the corresponding first processing module 51 is specifically used to determine the first conversation duration of the first user within a first preset time period based on the first voice data in response to the second state being a non-calling state.
[0193] In an alternative example, Figure 16 FIG. 5 is a schematic structural diagram of a fourth processing module 54 provided in an exemplary embodiment of the present disclosure. In this example, the fourth processing module 54 includes:
[0194] The fourth processing unit 541 is used to determine a first ratio based on the total conversation duration and the total duration of the first preset time period; the fifth processing unit 542 is used to determine a second ratio based on the first conversation duration and the total conversation duration; the sixth processing unit 543 is used to determine a third ratio based on the second conversation duration and the total conversation duration; the seventh processing unit 544 is used to determine whether the first conversation duration, the second conversation duration and the total conversation duration meet the preset conditions based on the first ratio, the second ratio and the third ratio.
[0195] Exemplary electronic devices
[0196] An embodiment of the present disclosure further provides an electronic device, comprising: a memory for storing a computer program;
[0197] The processor is configured to execute the computer program stored in the memory, and when the computer program is executed, the method for detecting the user conversation status described in any one of the above embodiments of the present disclosure is implemented.
[0198] Figure 17 FIG. 1 is a schematic diagram of a structure of an application embodiment of an electronic device disclosed in the present invention. In this embodiment, the electronic device 10 includes one or more processors 11 and a memory 12.
[0199] The processor 11 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.
[0200] The memory 12 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, a flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 11 may execute the program instructions to implement the methods of the various embodiments of the present disclosure described above and / or other desired functions. Various contents such as input signals, signal components, noise components, etc. may also be stored in the computer-readable storage medium.
[0201] In one example, the electronic device 10 may further include an input device 13 and an output device 14 , and these components are interconnected via a bus system and / or other forms of connection mechanisms (not shown).
[0202] For example, the input device 13 may be the aforementioned microphone or microphone array, used to capture input signals from a sound source.
[0203] In addition, the input device 13 may also include, for example, a keyboard, a mouse, and the like.
[0204] The output device 14 can output various information to the outside, including determined distance information, direction information, etc. The output device 14 can include, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and the like.
[0205] Of course, to simplify, Figure 17 Only some of the components related to the present disclosure in the electronic device 10 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, the electronic device 10 may further include any other appropriate components according to specific application scenarios.
[0206] Exemplary computer program products and computer-readable storage media
[0207] In addition to the above-mentioned methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions, which, when executed by a processor, enable the processor to perform the steps of the method according to various embodiments of the present disclosure described in the above-mentioned "Exemplary Method" section of this specification.
[0208] The computer program product may be written in any combination of one or more programming languages to implement the operations of the disclosed embodiments, including object-oriented programming languages such as Java, C++, and conventional procedural programming languages such as C or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0209] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, enable the processor to execute the steps of the method according to various embodiments of the present disclosure described in the above “Exemplary Method” section of this specification.
[0210] The computer-readable storage medium can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can, for example, include but is not limited to a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0211] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in this disclosure are merely illustrative and not restrictive, and should not be construed as necessarily possessed by each embodiment of the present disclosure. Furthermore, the specific details disclosed above are provided for illustrative purposes and to facilitate understanding, rather than as limitations. These details do not limit the present disclosure to necessarily being implemented using these specific details.
[0212] Each embodiment in this specification is described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. References to the same or similar parts between the various embodiments are sufficient. For system embodiments, since they largely correspond to method embodiments, their description is relatively simple. For relevant parts, references to the description of the method embodiments are sufficient.
[0213] The block diagrams of the devices, devices, equipment, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As will be appreciated by those skilled in the art, these devices, devices, equipment, and systems can be connected, arranged, or configured in any manner. Words such as "include," "comprise," "have," and the like are open-ended words, meaning "including but not limited to," and can be used interchangeably therewith. The words "or" and "and" used herein refer to the words "and / or" and can be used interchangeably therewith, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to," and can be used interchangeably therewith.
[0214] The methods and apparatus of the present disclosure may be implemented in many ways. For example, the methods and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above unless otherwise specified. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the methods according to the present disclosure. Thus, the present disclosure also covers recording media that store programs for executing the methods according to the present disclosure.
[0215] It should also be noted that in the apparatus, device, and method of the present disclosure, each component or each step can be decomposed and / or recombined. Such decomposition and / or recombination should be regarded as equivalent solutions of the present disclosure.
[0216] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0217] The above description has been provided for the purpose of illustration and description. In addition, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although a number of example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A method for detecting a user's conversation status, comprising: determining a first conversation duration of the first user within a first preset time period based on the first voice data; determining a second conversation duration of the second user within the first preset time period based on the second voice data; Determining a total conversation duration within the first preset time period based on the first voice data and the second voice data; Determining whether the first conversation duration, the second conversation duration, and the total conversation duration meet preset conditions based on preset rules; In response to the first conversation duration, the second conversation duration, and the total conversation duration satisfying a preset condition, determining that the first user and the second user are in a conversation state; The determining, based on the first voice data, of a first conversation duration of the first user within a first preset time period includes: determining the first conversation duration of the first user based on the first voice data and the corresponding first image data; and / or, The determining, based on the second voice data, a second conversation duration of the second user within the first preset time period includes: determining the second conversation duration of the second user based on the second voice data and the corresponding second image data; The determining, based on the first voice data and the corresponding first image data, the first conversation duration of the first user includes: determining, based on the first voice data and the first image data, a voice activation duration of the first user within the first preset time period; and determining, based on the voice activation duration of the first user within the first preset time period, a first conversation duration; The determining, based on the second voice data and the corresponding second image data, the second conversation duration of the second user includes: Based on the second voice data and the second image data, determine the voice activation duration of the second user within the first preset time period; based on the voice activation duration of the second user within the first preset time period, determine the second conversation duration.
2. The method according to claim 1, wherein The determining, based on the first voice data and the corresponding first image data, the first conversation duration of the first user includes: taking each speech frame in the first speech data as a current speech frame, taking an image frame in the first image data corresponding to the current speech frame as a current image frame, and determining a first speech activation state corresponding to the current frame based on the current speech frame and the current image frame; Determine whether the current cycle has ended; In response to the end of the current cycle, determining a first activation duration in the current cycle based on the first voice activation state corresponding to each frame in the current cycle, and entering processing of a next cycle until processing of the first preset time period is completed; determining the first conversation duration based on the first activation duration in each cycle within the first preset time period; In response to the current cycle not ending, recording the first voice activation state corresponding to the current frame, and entering the processing of the next frame until the current cycle ends.
3. The method according to claim 2, wherein: After determining, in response to the end of the current cycle, the first activation duration in the current cycle based on the first voice activation state corresponding to each frame in the current cycle, the method further includes: Adding the first activation duration in the current cycle to a first queue; The determining the first conversation duration based on the first activation duration in each cycle within the first preset time period includes: The first conversation duration is determined based on the first activation duration in each cycle in the first queue.
4. The method according to claim 2, wherein: After determining the first voice activation state corresponding to the current frame based on the current voice frame and the current image frame, the method further includes: In response to the first speech activation state corresponding to the current frame being active, determining a lip movement state corresponding to the current image frame based on a trained lip movement detection model; In response to the lip movement state corresponding to the current image frame being dynamic, determining that the first voice activation state corresponding to the current frame is active; In response to the lip movement state corresponding to the current image frame being static, it is determined that the first voice activation state corresponding to the current frame is inactive.
5. The method according to claim 1, wherein The determining, based on the second voice data and the corresponding second image data, the second conversation duration of the second user includes: taking each speech frame in the second speech data as a current speech frame, taking an image frame in the second image data corresponding to the current speech frame as a current image frame, and determining a second speech activation state corresponding to the current frame based on the current speech frame and the current image frame; Determine whether the current cycle has ended; In response to the end of the current cycle, determining a second activation duration in the current cycle based on the second voice activation state corresponding to each frame in the current cycle, and entering processing of a next cycle until processing of the first preset time period is completed; determining the second conversation duration based on the second activation duration in each cycle within the first preset time period; In response to the current cycle not ending, recording the second voice activation state corresponding to the current frame, and entering the processing of the next frame until the current cycle ends.
6. The method according to claim 5, wherein: After determining, in response to the end of the current cycle, the second activation duration in the current cycle based on the second voice activation state corresponding to each frame in the current cycle, the method further includes: Adding the second activation duration in the current cycle to a second queue; The determining the second conversation duration based on the second activation duration in each cycle within the first preset time period includes: The second conversation duration is determined based on the second activation duration in each cycle in the second queue.
7. The method according to claim 1, wherein The determining, based on the first voice data and the corresponding first image data, the first conversation duration of the first user includes: determining a first fusion feature based on the first speech data and the first image data; Detecting the first fusion feature based on a pre-trained multimodal speech endpoint detection model to obtain a first detection result, where the first detection result includes first probability values corresponding to a plurality of time points within the first preset time period; Determining the first conversation duration of the first user based on the first detection result; The determining, based on the second voice data and the corresponding second image data, the second conversation duration of the second user includes: determining a second fusion feature based on the second speech data and the second image data; Detecting the second fusion feature based on the pre-trained multimodal speech endpoint detection model to obtain a second detection result, where the second detection result includes second probability values corresponding to multiple time points within the first preset time period; The second conversation duration of the second user is determined based on the second detection result.
8. The method according to claim 7, wherein: The determining, based on the first detection result, the first conversation duration of the first user includes: Determining a first time point in a conversation state among the time points based on first probability values corresponding to the plurality of time points within the first preset time period; determining the first conversation duration based on the first time point in the conversation state among the time points; The determining, based on the second detection result, the second conversation duration of the second user includes: Determining a second time point in the conversation state among the time points based on second probability values corresponding to the plurality of time points within the first preset time period; The second conversation duration is determined based on a second time point in the conversation state among the time points.
9. The method according to claim 1, wherein Before determining the first conversation duration of the first user within the first preset time period based on the first voice data, the method further includes: determining a first state of the first user based on first image data corresponding to the first voice data; In response to the first state being a non-calling state, determining a second state of the second user based on second image data corresponding to the second voice data; The determining, based on the first voice data, a first conversation duration of the first user within a first preset time period includes: In response to the second state being a non-calling state, a first conversation duration of the first user within a first preset time period is determined based on the first voice data.
10. The method according to any one of claims 1 to 9, wherein: The determining, based on the first voice data and the second voice data, a total conversation duration within the first preset time period includes: The total conversation duration within the first preset time period is determined based on the first voice data, the first image data corresponding to the first voice data, the second voice data, and the second image data corresponding to the second voice data.
11. The method according to any one of claims 1 to 9, wherein: The determining, based on a preset rule, whether the first conversation duration, the second conversation duration, and the total conversation duration meet a preset condition includes: determining a first ratio based on the total conversation duration and the total duration of the first preset time period; determining a second ratio based on the first conversation duration and the total conversation duration; determining a third ratio based on the second conversation duration and the total conversation duration; Based on the first ratio, the second ratio, and the third ratio, it is determined whether the first conversation duration, the second conversation duration, and the total conversation duration meet a preset condition.
12. A device for detecting a user's conversation status, comprising: A first processing module, configured to determine a first conversation duration of a first user within a first preset time period based on the first voice data; a second processing module, configured to determine a second conversation duration of the second user within the first preset time period based on the second voice data; a third processing module, configured to determine a total conversation duration within the first preset time period based on the first voice data and the second voice data; a fourth processing module, configured to determine, based on a preset rule, whether the first conversation duration, the second conversation duration, and the total conversation duration meet a preset condition; a fifth processing module, configured to determine that the first user and the second user are in a conversation state in response to the first conversation duration, the second conversation duration, and the total conversation duration satisfying a preset condition; Wherein, the first processing module includes: a first processing unit, configured to determine the first conversation duration of the first user based on the first voice data and the corresponding first image data; The second processing module includes: a second processing unit, configured to determine the second conversation duration of the second user based on the second voice data and the corresponding second image data; The first processing unit is specifically configured to: determine, based on the first voice data and the first image data, a voice activation duration of the first user within the first preset time period; and determine, based on the voice activation duration of the first user within the first preset time period, a first conversation duration; The second processing unit is specifically used to: determine the voice activation duration of the second user within the first preset time period based on the second voice data and the second image data; and determine the second conversation duration based on the voice activation duration of the second user within the first preset time period.
13. A computer-readable storage medium storing a computer program, wherein the computer program is used to execute the method for detecting the user conversation status according to any one of claims 1 to 11.
14. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method for detecting the user conversation status as described in any one of claims 1-11.
Citation Information
Patent Citations
Teaching mode analysis method and system
CN112599135A