Real-time video dialogue method and device of electronic equipment

By turning on the screen freezing function in real-time video conversations to lock the preview screen, the problem of users needing to continuously track moving or disappearing the target object is solved, and the stability and accuracy of the conversation are achieved.

CN120034613APending Publication Date: 2025-05-23SAMSUNG GUANGZHOU MOBILE R&D CENT +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510113333.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

During the live video conversation between users and artificial intelligence assistants, target objects that are easily moved or disappeared cause users to continuously adjust the camera of electronic devices, which affects the consistency and accuracy of the conversation.

Method used

By turning on the screen freezing function in the real-time video conversation, lock the current preview screen, and conduct real-time video conversation based on the locked preview screen until the screen freezing function is turned off.

Benefits of technology

This enables users to conduct conversations without continuously tracking the target object, avoids incorrect answering questions caused by the target object moving or disappearing, and improves the stability and accuracy of the conversation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034613A_ABST
    Figure CN120034613A_ABST
Patent Text Reader

Abstract

The invention provides a real-time video conversation method and device of electronic equipment. The real-time video dialogue method of the electronic equipment comprises the following steps: displaying an interface for a real-time video dialogue between a user and an artificial intelligence assistant of the electronic equipment in response to a preset operation; determining a locked preview picture of the real-time video dialogue in response to the fact that a picture freeze-frame function of the real-time video dialogue is started; the locked preview picture is displayed in the interface in a specific mode; and performing the real-time video dialogue based on the locked preview picture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the field of electronic technology, and more specifically, to a real-time video conversation method and device for an electronic device. Background Art

[0002] With the rapid development of Internet technology, Real-Time Communication (RTC) technology has been widely used in various fields, especially in the field of voice and video conversation. RTC technology has brought users the ultimate communication experience. On this basis, RTC intelligent voice / video conversation integrated with artificial intelligence technology has gradually become the representative of the new generation of communication methods.

[0003] Through RTC intelligent voice / video technology, people can directly have real-time video conversations with artificial intelligence assistants, including but not limited to chatting, problem consultation, video object recognition and other scenarios, providing a variety of services for different groups of people. Summary of the invention

[0004] An exemplary embodiment of the present disclosure is to provide a real-time video conversation method and apparatus for an electronic device, which facilitates a user to continue to have a conversation with an artificial intelligence assistant regarding a target object (e.g., an object whose information the user expects to learn from the artificial intelligence assistant) during a real-time video conversation between the user and the artificial intelligence assistant.

[0005] According to a first aspect of an embodiment of the present disclosure, a real-time video conversation method for an electronic device is provided, comprising: in response to a preset operation, displaying an interface for a user to conduct a real-time video conversation with an artificial intelligence assistant of the electronic device; in response to a freeze frame function of the real-time video conversation being turned on, determining a locked preview screen of the real-time video conversation; displaying the locked preview screen in a specific manner in the interface; and conducting the real-time video conversation based on the locked preview screen.

[0006] Optionally, the method further includes: in response to a freeze frame function of the real-time video conversation being turned off, displaying a real-time preview screen of the real-time video conversation in the interface; and conducting the real-time video conversation based on the real-time preview screen.

[0007] Optionally, in response to the freeze frame function of the real-time video conversation being turned on, the step of determining the locked preview screen of the real-time video conversation includes: in response to the freeze frame function of the real-time video conversation being turned on, taking a screenshot of the current preview screen of the real-time video conversation, and using the screenshot as the locked preview screen; or, in response to the freeze frame function of the real-time video conversation being turned on, controlling the camera of the electronic device to capture an image, and using the captured image as the locked preview screen; or, in response to the freeze frame function of the real-time video conversation being turned on, providing candidate preview screens to the user, and using the preview screen selected by the user from the candidate preview screens as the locked preview screen, wherein the candidate preview screens include the current preview screen and a predetermined number of historical preview screens adjacent to the current preview screen.

[0008] Optionally, the step of displaying the locked preview screen in a specific manner in the interface includes: in the interface, using the locked preview screen to replace the real-time preview screen of the real-time video conversation for display; or, in the interface, superimposing the locked preview screen on the real-time preview screen of the real-time video conversation.

[0009] Optionally, the step of conducting the real-time video conversation based on the locked preview screen includes: sending the locked preview screen to the server at preset time intervals; sending audio data collected in real time by the electronic device to the server; obtaining reply data from the server by the server calling the machine learning model based on the locked preview screen and the audio data collected in real time; and providing a reply result to the user based on the reply data.

[0010] Optionally, the step of conducting the real-time video conversation based on the locked preview screen includes: sending the locked preview screen to the server, and request information for requesting the server to store and continue to use the locked preview screen provided this time; sending audio data collected in real time by the electronic device to the server; obtaining from the server reply data obtained by the server calling a machine learning model based on the locked preview screen and the audio data collected in real time; and providing a reply result to the user based on the reply data.

[0011] Optionally, the step of conducting the real-time video conversation based on the real-time preview screen includes: sending the real-time preview screen and the audio data collected in real time by the electronic device to a server; obtaining from the server reply data obtained by the server calling a machine learning model based on the real-time preview screen and the audio data collected in real time; and providing a reply result to the user based on the reply data.

[0012] Optionally, the step of conducting the real-time video conversation based on the locked preview screen includes: calling a machine learning model to obtain reply data based on the locked preview screen and the audio data collected by the electronic device in real time; and providing a reply result to the user based on the reply data.

[0013] Optionally, the step of conducting the real-time video conversation based on the real-time preview screen includes: calling a machine learning model to obtain reply data based on the real-time preview screen and the audio data collected in real time by the electronic device using the machine learning model; and providing a reply result to the user based on the reply data.

[0014] Optionally, it also includes: in response to a drawing operation performed by a user on the locked preview screen, superimposing and displaying drawing traces formed by the drawing operation on the locked preview screen; wherein the step of conducting the real-time video conversation based on the locked preview screen includes: conducting the real-time video conversation based on the locked preview screen superimposed with the drawing traces.

[0015] Optionally, it also includes: in response to the user's drawing operation on the real-time preview screen, superimposing and displaying drawing traces formed by the drawing operation on the real-time preview screen; wherein the step of conducting the real-time video conversation based on the real-time preview screen includes: conducting the real-time video conversation based on the real-time preview screen superimposed with the drawing traces.

[0016] According to a second aspect of an embodiment of the present disclosure, a real-time video conversation device of an electronic device is provided, comprising: an interface display unit, configured to display an interface for a user to conduct a real-time video conversation with an artificial intelligence assistant of the electronic device in response to a preset operation; a screen locking unit, configured to determine a locked preview screen of the real-time video conversation in response to a screen freeze function of the real-time video conversation being turned on; a video conversation unit, configured to conduct the real-time video conversation based on the locked preview screen in response to the screen freeze function of the real-time video conversation being turned on; wherein the interface display unit is further configured to display the locked preview screen in a specific manner in the interface in response to the screen freeze function of the real-time video conversation being turned on.

[0017] According to a third aspect of an embodiment of the present disclosure, a computer-readable storage medium storing a computer program is provided, and when the computer program is executed by a processor, the real-time video conversation method of the electronic device as described above is implemented.

[0018] According to a fourth aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory storing a computer program, wherein when the computer program is executed by the processor, the real-time video conversation method of the electronic device as described above is implemented.

[0019] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising computer executable instructions, which, when executed by a processor, can implement the real-time video conversation method of an electronic device as described above.

[0020] According to the real-time video conversation method and device of an electronic device of an exemplary embodiment of the present disclosure, during a real-time video conversation between a user and an artificial intelligence assistant, when faced with a target object that is easy to move or disappear, the user can better lock the target object so as to have a conversation with the artificial intelligence assistant around the target object without having to continuously aim the electronic device at the target object and follow the target object to shoot, which can also avoid the problem of the target object disappearing.

[0021] Additional aspects and / or advantages of the present general inventive concept will be set forth in part in the following description and in part will be apparent from the description or may be learned through practice of the present general inventive concept. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The above and other objects and features of the exemplary embodiments of the present disclosure will become more apparent through the following description in conjunction with the accompanying drawings which exemplarily illustrate the embodiments, in which: Figure 1 A flowchart showing a real-time video conversation method of an electronic device according to an exemplary embodiment of the present disclosure; Figure 2 An example of an interface showing a real-time video conversation between a user and an artificial intelligence assistant according to an exemplary embodiment of the present disclosure is shown; Figure 3 An example of turning on / off a freeze frame function according to an exemplary embodiment of the present disclosure is shown; Figure 4 A flowchart showing a method for conducting a real-time video conversation based on a locked preview screen according to a first exemplary embodiment of the present disclosure; Figure 5 A flowchart showing a method for conducting a real-time video conversation based on a locked preview screen according to a second exemplary embodiment of the present disclosure; Figure 6 A flowchart showing a method for conducting a real-time video conversation based on a locked preview screen according to a third exemplary embodiment of the present disclosure; Figure 7An example of a method for conducting a real-time video conversation based on a locked preview screen according to a first exemplary embodiment of the present disclosure is shown; Figure 8 An example of a method for conducting a real-time video conversation based on a locked preview screen according to a second exemplary embodiment of the present disclosure is shown; Fig. 9 A flowchart showing a method for conducting a real-time video conversation based on a real-time preview screen according to a first exemplary embodiment of the present disclosure; Fig.10 A flowchart showing a method for conducting a real-time video conversation based on a real-time preview screen according to a second exemplary embodiment of the present disclosure; Fig.11 An example of a method for conducting a real-time video conversation based on a real-time preview screen according to a first exemplary embodiment of the present disclosure is shown; Fig.12 An example of using a brush function according to an exemplary embodiment of the present disclosure is shown; Fig.13 A structural block diagram of a real-time video conversation apparatus of an electronic device according to an exemplary embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0023] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings, wherein like reference numerals refer to like parts throughout. The embodiments will be described below with reference to the drawings in order to explain the present disclosure.

[0024] During the real-time video conversation between the user and the AI ​​assistant, for a target object that is easy to move, the user needs to keep the camera of the electronic device aimed at the target object and follow the target object to shoot, so as to avoid the AI ​​assistant being unable to correctly answer the user's inquiries about the target object because the target object is not in the lens. In addition, for a target object that is easy to disappear, the following situation is likely to occur: before the user completes the inquiry about the target object, the target object has disappeared.

[0025] For example, scenario 1: The user turns on the RTC video conversation function (RTC Video) of the AI ​​assistant and wants to know the information of a kitten. Since kittens are lively and active, the user needs to hold up the electronic device to track and shoot the kitten along the kitten's movement route, and ask the AI ​​assistant "What breed of kitten is this?" If the kitten is not tracked and shot, the following situation may occur: the kitten leaves, and only a chair is left in the video. The AI ​​assistant will answer "This is not a kitten, it is a bench", instead of "This is a Chinese rural cat". Scenario 2: The user turns on the RTC video conversation function of the AI ​​assistant and wants to know the information of the car in the advertisement played on the display screen on the exterior wall of the shopping mall, so he holds up the electronic device to shoot the display screen and asks the AI ​​assistant "What brand of car is this and what model is it?" However, during the inquiry, the screen content of the exterior display screen changes and starts playing the next cosmetics advertisement. The AI ​​assistant can only reply "This is a Chinese rural cat" based on the content of the next advertisement. Makeup, not cars."

[0026] Considering the above-mentioned situation in the prior art, the present disclosure proposes that: during a real-time video conversation between a user and an artificial intelligence assistant, if the user wishes to inquire about an object (i.e., a target object) in the current video preview screen, the freeze function can be activated to lock the current video preview screen, and a real-time video conversation can be conducted based on the locked preview screen until the freeze function is turned off. This eliminates the need for the user to continuously aim the electronic device at the target object and follow the target object to shoot, and can effectively avoid obtaining an incorrect answer due to the movement or disappearance of the target object.

[0027] According to the present disclosure, for example, scenario 1: the user turns on the RTC video conversation function of the artificial intelligence assistant and wants to know the information of a kitten. When the kitten appears in the current preview screen, the user can immediately start the freeze frame function to lock the preview screen containing the kitten, and then ask the artificial intelligence assistant "What breed of kitten is this?" Even if the kitten leaves after the preview screen is locked, the artificial intelligence assistant can still accurately answer "This is a Chinese rural cat" based on the locked preview screen. In addition, the user can continue to ask other questions based on the locked preview screen without considering the movement of the kitten. Scenario 2: The user turns on the RTC video conversation function of the artificial intelligence assistant and wants to know the information of the car in the advertisement played on the display screen on the exterior wall of the shopping mall. When the car advertisement is played on the display screen on the exterior wall, the user can shoot the car advertisement and immediately start the freeze frame function to lock the preview screen containing the car, and then ask the artificial intelligence assistant "What brand of car is this car and what specific model is it?" Even if the next advertisement is played on the display screen on the exterior wall after the preview screen is locked, the artificial intelligence assistant can still accurately answer "This is a Chinese rural cat" based on the locked preview screen. Brand of car, specifically Model", in addition, the user can continue to ask other questions based on the locked preview screen (for example, what is the specific price?) without considering whether the advertisement for the car is no longer playing.

[0028] The following will be combined Figures 1 to 13 Exemplary embodiments of the present disclosure are described in detail.

[0029] Figure 1 A flow chart showing a real-time video conversation method of an electronic device according to an exemplary embodiment of the present disclosure is shown.

[0030] The real-time video conversation method can be implemented by a computer program. As an example, the real-time video conversation can be performed by an application (e.g., an artificial intelligence assistant APP) installed in an electronic device, or by a functional program implemented in an operating system of the electronic device.

[0031] As an example, the electronic device may be a mobile communication terminal (eg, a smart phone), a smart wearable device (eg, a smart watch), a game console, a digital multimedia player, or other electronic device.

[0032] Reference Figure 1 In step S100, in response to a preset operation, an interface for a user to conduct a real-time video conversation with an artificial intelligence assistant of an electronic device (hereinafter, also referred to as a real-time video conversation interface) is displayed.

[0033] As an exemplary embodiment, step S100 may include: in response to a user operation for waking up the artificial intelligence assistant, displaying the main interface of the artificial intelligence assistant; then, in response to a user operation for having a real-time conversation with the artificial intelligence assistant, displaying an interface for the user to have a real-time conversation with the artificial intelligence assistant; next, in response to a user operation for having a real-time video conversation with the artificial intelligence assistant, enabling the camera and displaying an interface for the user to have a real-time video conversation with the artificial intelligence assistant. In the real-time video conversation interface, the video conversation screen content (i.e., the camera preview screen) will be displayed, and the user can have a conversation with the artificial intelligence assistant based on the preview screen. For example, the artificial intelligence assistant can be asked questions based on the preview screen, and the artificial intelligence assistant will answer based on the preview screen and the user's questions.

[0034] As an example, Figure 2 As shown, in response to a user operation for waking up an artificial intelligence assistant (such as Bixby), the Bixby interface is displayed; then, in response to a user operation of clicking the RTC dialogue control in the Bixby interface, the RTC chat interface of Bixby is displayed; next, in response to a user operation of clicking the RTC Video control in the RTC chat interface, the RTCVideo dialogue function is turned on, specifically, the camera of the electronic device is enabled, and the RTC Video dialogue interface between the user and Bixby is displayed.

[0035] It should be understood that the above-mentioned preset operation is not limited to manual operation, but may also be other types of operations, such as voice command operation.

[0036] In step S200, in response to the freeze function of the real-time video conversation being turned on, a locked preview picture of the real-time video conversation is determined.

[0037] As an example, Figure 3 As shown, the user can lock the current preview screen through the screen freeze control displayed in the real-time video conversation interface. In addition, other operation methods can also be used to turn on / off the screen freeze function, such as voice command operation.

[0038] As an exemplary embodiment, in response to the freeze function of the real-time video conversation being turned on, a screenshot of the current preview screen of the real-time video conversation is taken, and the obtained screenshot is used as the locked preview screen. Specifically, in response to the freeze function of the real-time video conversation being turned on, the camera preview screen currently displayed in the real-time video conversation interface is obtained by taking a screenshot to obtain the locked preview screen.

[0039] As another exemplary embodiment, in response to the freezing function of the real-time video conversation being turned on, the camera of the electronic device is controlled to capture an image, and the captured image is used as the locked preview image.

[0040] In the above two exemplary embodiments, the current preview screen of the real-time video conversation is directly locked. In addition, as another exemplary embodiment, in response to the freeze function of the real-time video conversation being turned on, a candidate preview screen is provided to the user, and the preview screen selected by the user from the candidate preview screen is used as the locked preview screen. Among them, the candidate preview screen includes the current preview screen and a predetermined number of historical preview screens adjacent to the current preview screen. As an example, the method of providing the user with the candidate preview screen may include but is not limited to: displaying the current preview screen in the form of a large picture, and displaying a predetermined number of historical preview screens adjacent to the current preview screen in the form of a small picture. According to an exemplary embodiment of the present disclosure, even if the current preview screen does not include a target object that is easy to move or disappear, a historical preview screen including the target object can be selected as the locked preview screen.

[0041] In step S300, the locked preview screen is displayed in a specific manner in the real-time video conversation interface.

[0042] As an exemplary embodiment, in the real-time video conversation interface, a locked preview screen is used to replace the real-time preview screen of the real-time video conversation for display. Specifically, if the picture freeze function is not turned on, the real-time preview screen is displayed in the real-time video conversation interface; if the picture freeze function is turned on, the real-time preview screen is not displayed in the real-time video conversation interface, but the locked preview screen is displayed.

[0043] As another exemplary embodiment, in the real-time video conversation interface, a locked preview screen is superimposed and displayed on the real-time preview screen of the real-time video conversation. Specifically, if the picture freeze function is not turned on, the real-time preview screen is displayed in the real-time video conversation interface; if the picture freeze function is turned on, not only the real-time preview screen is displayed in the real-time video conversation interface, but also the locked preview screen is superimposed and displayed on the real-time preview screen. It should be understood that the present disclosure does not limit the specific form of superimposed display. For example, the locked preview screen can be superimposed and displayed on a partial area (e.g., the center area or the upper right area, etc.) of the real-time preview screen.

[0044] It should be understood that the real-time preview picture is the latest video frame in the video continuously shot by the camera in real time. Therefore, the real-time preview picture is not fixed but changes with time.

[0045] In step S400, a real-time video conversation is conducted based on the locked preview screen.

[0046] Specifically, when the freeze frame function is turned on, the real-time video conversation always revolves around the locked preview screen. Figures 4 to 8An exemplary embodiment of step S400 is described in detail and will not be expanded here.

[0047] In addition, the real-time video conversation method of an electronic device according to an exemplary embodiment of the present disclosure may also include: in response to the freeze function of the real-time video conversation being turned off, displaying a real-time preview screen of the real-time video conversation in the real-time video conversation interface; and conducting a real-time video conversation based on the real-time preview screen.

[0048] As an example, Figure 3 As shown, the user can unlock the preview screen through the freeze control displayed in the real-time video conversation interface, and restore the real-time preview screen to be displayed in the real-time video conversation interface instead of the locked preview screen; and conduct a real-time video conversation based on the real-time preview screen.

[0049] In addition, the real-time video conversation method of an electronic device according to an exemplary embodiment of the present disclosure may also include: displaying a real-time preview screen of the real-time video conversation in the real-time video conversation interface before the freeze frame function of the real-time video conversation is turned on; and conducting a real-time video conversation based on the real-time preview screen.

[0050] In other words, only when the freeze function is turned on, the real-time video conversation is based on the locked preview screen. When the freeze function is turned off, the real-time video conversation is based on the real-time preview screen. Figures 9 to 11 An exemplary embodiment of a method for conducting a real-time video conversation based on a real-time preview screen is described in detail, which will not be expanded here.

[0051] In addition, according to an exemplary embodiment of the present disclosure, a brush function is also provided for the real-time video conversation interface. As an exemplary embodiment, when the freeze function is in a turned off state, in response to the user's painting operation on the real-time preview screen displayed in the real-time video conversation interface, the painting traces formed by the painting operation are superimposed on the real-time preview screen; accordingly, a real-time video conversation is conducted based on the real-time preview screen with the superimposed painting traces.

[0052] As another exemplary embodiment, when the freeze function is turned on, in response to a user's drawing operation on a locked preview screen displayed in the real-time video conversation interface, a drawing trace formed by the drawing operation is superimposed on the locked preview screen; accordingly, a real-time video conversation is conducted based on the locked preview screen superimposed with the drawing trace. Fig.12As shown in the figure, in the scene where the picture is frozen, the user can use the brush function to doodle or circle the current frozen picture (that is, the locked preview picture). Every time the user uses the brush function to draw on the locked preview picture, a real-time video conversation will be carried out based on the latest picture with the trace of the drawing operation.

[0053] According to the brush function provided by the present disclosure, it is possible to focus on the details of the preview image (for example, a certain part of the target object) to conduct a real-time video conversation so that the user can obtain the required information.

[0054] Figure 4 A flowchart of a method for conducting a real-time video conversation based on a locked preview screen according to a first exemplary embodiment of the present disclosure is shown. In this embodiment, a machine learning model is deployed in a server.

[0055] Reference Figure 4 In step S401, a locked preview screen is sent to the server every preset time period.

[0056] In step S402, the audio data collected by the electronic device in real time is sent to the server.

[0057] As an exemplary embodiment, the specific value of the preset duration can be determined according to actual conditions and specific needs, and the specific value of the preset duration can be fixed or dynamically changing. For example, the specific value of the preset duration can be determined based on the frequency of sending audio data to the server.

[0058] In step S403, response data is obtained from the server by the server calling the machine learning model based on the locked preview screen and the audio data collected in real time.

[0059] In step S404, a reply result is provided to the user based on the reply data.

[0060] As an example, the server may first process the locked preview screen and the audio data collected in real time (e.g., voice to text), and then input the processed data into the machine learning model to obtain the output data of the machine learning model. As another example, the server may directly input the locked preview screen and the audio data collected in real time into the machine learning model to obtain the output data of the machine learning model. In addition, as an example, the server may process the output data of the machine learning model (e.g., text to speech) to obtain reply data. As another example, the server may directly use the output data of the machine learning model as reply data.

[0061] In addition, as an example, when the server detects that the user has spoken in the voice data and the speaking has been completed, it can call the machine learning model to combine the information in the voice data and the information in the preview screen to infer the answer to the question.

[0062] As an exemplary embodiment, the reply result may include but is not limited to: a voice reply result. As an example, the reply data may be processed (e.g., text-to-speech) to obtain the reply result. As another example, the reply data (e.g., when the reply data is voice data) may be directly used as the reply result.

[0063] As an exemplary embodiment, the machine learning model (also referred to as an artificial intelligence model) used in the present disclosure may be a large multimodal model for image recognition, which is capable of answering user questions based on images.

[0064] The machine learning model can be obtained through training. Here, “obtained through training” refers to a machine learning model that is trained based on multiple training data through a training algorithm.

[0065] As an example, a machine learning model may include multiple neural network layers. Each of the multiple neural network layers includes multiple weight values, and the neural network calculation is performed by calculating between the calculation result of the previous layer and the multiple weight values. It should be understood that the present disclosure does not limit the form of specific input data, the form of output data, and the specific model structure of the machine learning model.

[0066] Figure 7 An example of a method of conducting a real-time video conversation based on a locked preview screen according to a first exemplary embodiment of the present disclosure is shown.

[0067] like Figure 7 As shown, in step S701, currently collected voice data and a locked preview image (ie, video frame S) are obtained.

[0068] In step S702, the video frame S is cached so that the video frame S is subsequently sent to the server at a preset time interval.

[0069] In step S703, the acquired voice data and video frame S are sent to the server.

[0070] In step S704, the server processes the received voice data and video frame S, inputs the processed data into the machine learning model, and obtains voice response data based on the output data of the machine learning model.

[0071] In step S705, voice reply data is obtained from the server and played to the user.

[0072] In step S706, newly collected voice data is obtained, and the newly obtained voice data and the cached video frame S are sent to the server. The server processes the newly received voice data and video frame S, inputs the processed data into the machine learning model, and obtains new voice reply data based on the output data of the machine learning model. The electronic device obtains and plays the new voice reply data from the server. It should be understood that step S706 can be performed periodically before turning off the freeze function.

[0073] Figure 5 A flowchart of a method for conducting a real-time video conversation based on a locked preview screen according to a second exemplary embodiment of the present disclosure is shown. In this embodiment, a machine learning model is deployed in a server.

[0074] Reference Figure 5 In step S501, the locked preview screen is sent to the server, as well as request information for requesting the server to store and continue to use the locked preview screen provided this time.

[0075] In step S502, the audio data collected in real time by the electronic device is sent to the server.

[0076] In step S503, response data is obtained from the server by the server calling the machine learning model based on the locked preview screen and the audio data collected in real time.

[0077] In step S504, a reply result is provided to the user based on the reply data.

[0078] Figure 8 An example of a method of conducting a real-time video conversation based on a locked preview screen according to a second exemplary embodiment of the present disclosure is shown.

[0079] like Figure 8 As shown, in step S801, currently collected voice data and a locked preview image (ie, video frame S) are obtained.

[0080] In step S802, the acquired voice data and video frame S (carrying request information for requesting the server to store and continuously use the video frame S provided this time) are sent to the server.

[0081] In step S803, after receiving the video frame S carrying the request information, the server caches the video frame S.

[0082] In step S804, the server processes the received voice data and video frame S, inputs the processed data into the machine learning model, and obtains voice response data based on the output data of the machine learning model.

[0083] In step S805, voice reply data is obtained from the server and played to the user.

[0084] In step S806, the electronic device only obtains and sends the newly collected voice data to the server, the server processes the newly received voice data and the cached video frame S, inputs the processed data into the machine learning model, and obtains new voice reply data based on the output data of the machine learning model, and the electronic device obtains and plays the new voice reply data from the server. It should be understood that step S806 can be performed periodically before turning off the freeze function.

[0085] Since the video frame S sent in step S802 carries the request information, there is no need to send the video frame S to the server multiple times. The video frame S carrying the request information only needs to be sent once, so that the server can continue to use the video frame S until the freeze function is canceled.

[0086] It should be understood that before turning off the freeze frame function, if the video frame S changes, for example, the user draws on the locked preview screen (i.e., video frame S) displayed in the real-time video conversation interface, then the changed video frame S (i.e., the video frame S with superimposed painting traces and carrying request information for requesting the server to store and continue to use the video frame S provided this time) needs to be sent to the server.

[0087] In addition, in response to the freezing function being turned off, a request message is sent to the server to request not to continue using the video frame S, so as to restore the normal process.

[0088] Figure 6 A flowchart of a method for conducting a real-time video conversation based on a locked preview screen according to a third exemplary embodiment of the present disclosure is shown. In this embodiment, a machine learning model is deployed in an electronic device.

[0089] Reference Figure 6 In step S601, a machine learning model is called to obtain response data based on the locked preview screen and the audio data collected in real time by the electronic device.

[0090] In step S602, a reply result is provided to the user based on the reply data.

[0091] As an example, the electronic device may first process the locked preview screen and the audio data collected in real time (e.g., voice to text), and then input the processed data into the machine learning model to obtain the output data of the machine learning model. As another example, the electronic device may directly input the locked preview screen and the audio data collected in real time into the machine learning model to obtain the output data of the machine learning model. In addition, as an example, the electronic device may process the output data of the machine learning model (e.g., text to speech) to obtain reply data. As another example, the electronic device may directly use the output data of the machine learning model as reply data.

[0092] In addition, as an example, when the electronic device detects that the user has spoken in the voice data and the speaking has been completed, it can call the machine learning model to combine the information in the voice data and the information in the preview screen to infer the answer to the question.

[0093] As an exemplary embodiment, the reply result may include but is not limited to: a voice reply result. As an example, the reply data may be processed (e.g., text-to-speech) to obtain the reply result. As another example, the reply data (e.g., when the reply data is voice data) may be directly used as the reply result.

[0094] Fig. 9 A flowchart of a method for conducting a real-time video conversation based on a real-time preview screen according to a first exemplary embodiment of the present disclosure is shown. In this embodiment, a machine learning model is deployed in a server.

[0095] Reference Fig. 9 In step S901, a real-time preview image and audio data collected in real time by the electronic device are sent to the server.

[0096] In step S902, response data is obtained from the server by the server calling the machine learning model based on the real-time preview screen and the real-time collected audio data.

[0097] In step S903, a reply result is provided to the user based on the reply data.

[0098] Fig.11 An example of a method for conducting a real-time video conversation based on a real-time preview screen according to an exemplary embodiment of the present disclosure is shown.

[0099] like Fig.11 As shown, in step S1101, the currently collected voice data and the current video preview screen (ie, Fig.11 frames in the video).

[0100] In step S1102, the acquired voice data and video preview screen are sent to the server.

[0101] As an exemplary embodiment, in order to ensure that the server can obtain a real-time video preview image, the video preview image can be obtained from the camera hardware abstraction layer Camera HAL at a preset time, that is, a fixed time interval (e.g., every 0.25 seconds), and sent to the server. It should be understood that the preset time is not limited and can be set according to actual conditions and specific needs.

[0102] It should be understood that the frequency of acquiring voice data may be the same as or different from the frequency of acquiring video preview images, and the frequency of sending voice data to the server may be the same as or different from the frequency of sending video preview images to the server.

[0103] In step S1103, the server processes the received voice data and video preview screen, inputs the processed data into the machine learning model, and obtains voice reply data based on the output data of the machine learning model.

[0104] In step S1104, voice reply data is obtained from the server and played to the user.

[0105] In step S1105, newly collected voice data and new video preview screen are obtained, and the newly obtained voice data and video preview screen are sent to the server. The server processes the newly received voice data and video preview screen, inputs the processed data into the machine learning model, and obtains new voice reply data based on the output data of the machine learning model. The electronic device obtains and plays the new voice reply data from the server. It should be understood that step S1105 can be performed periodically before the freeze function is turned on.

[0106] Fig.10 A flowchart of a method for conducting a real-time video conversation based on a real-time preview screen according to a second exemplary embodiment of the present disclosure is shown. In this embodiment, a machine learning model is deployed in an electronic device.

[0107] Reference Fig.10 In step S1001, a machine learning model is called to obtain response data based on the real-time preview screen and the audio data collected in real time by the electronic device using the machine learning model.

[0108] In step S1002, a reply result is provided to the user based on the reply data.

[0109] Fig.13 A structural block diagram of a real-time video conversation apparatus of an electronic device according to an exemplary embodiment of the present disclosure is shown.

[0110] Reference Fig.13According to an exemplary embodiment of the present disclosure, a real-time video conversation apparatus of an electronic device includes: an interface display unit 100 , a screen locking unit 200 , and a video conversation unit 300 .

[0111] Specifically, the interface display unit 100 is configured to display an interface for a real-time video conversation between a user and an artificial intelligence assistant of the electronic device in response to a preset operation.

[0112] The screen locking unit 200 is configured to determine a locked preview screen of the real-time video conversation in response to the screen freeze function of the real-time video conversation being turned on.

[0113] The video conversation unit 300 is configured to conduct the real-time video conversation based on the locked preview picture in response to the picture freezing function of the real-time video conversation being turned on.

[0114] The interface display unit 100 is further configured to display the locked preview screen in a specific manner in the interface in response to the freezing function of the real-time video conversation being turned on.

[0115] As an exemplary embodiment, the interface display unit 100 may also be configured to: in response to the freeze frame function of the real-time video conversation being turned off, display a real-time preview picture of the real-time video conversation in the interface; wherein, the video conversation unit 300 may also be configured to: in response to the freeze frame function of the real-time video conversation being turned off, conduct the real-time video conversation based on the real-time preview picture.

[0116] As an exemplary embodiment, the screen locking unit 200 may be configured to: in response to the freeze frame function of the real-time video conversation being turned on, take a screenshot of the current preview screen of the real-time video conversation, and use the screenshot as the locked preview screen; or, in response to the freeze frame function of the real-time video conversation being turned on, control the camera of the electronic device to capture an image, and use the captured image as the locked preview screen; or, in response to the freeze frame function of the real-time video conversation being turned on, provide candidate preview screens to the user, and use the preview screen selected by the user from the candidate preview screens as the locked preview screen, wherein the candidate preview screens include the current preview screen and a predetermined number of historical preview screens adjacent to the current preview screen.

[0117] As an exemplary embodiment, the interface display unit 100 may be configured to: in the interface, use the locked preview screen to replace the real-time preview screen of the real-time video conversation for display; or, in the interface, superimpose the locked preview screen on the real-time preview screen of the real-time video conversation.

[0118] As an exemplary embodiment, the video conversation unit 300 may be configured to: send the locked preview screen to the server at preset time intervals; send the audio data collected in real time by the electronic device to the server; obtain from the server reply data obtained by the server calling the machine learning model based on the locked preview screen and the audio data collected in real time; and provide a reply result to the user based on the reply data.

[0119] As an exemplary embodiment, the video conversation unit 300 may be configured to: send the locked preview screen to the server, and request information for requesting the server to store and continue to use the locked preview screen provided this time; send the audio data collected in real time by the electronic device to the server; obtain from the server reply data obtained by the server calling a machine learning model based on the locked preview screen and the audio data collected in real time; and provide a reply result to the user based on the reply data.

[0120] As an exemplary embodiment, the video conversation unit 300 may be configured to: send the real-time preview screen and the audio data collected in real time by the electronic device to the server; obtain from the server reply data obtained by the server calling the machine learning model based on the real-time preview screen and the audio data collected in real time; and provide a reply result to the user based on the reply data.

[0121] As an exemplary embodiment, the video conversation unit 300 may be configured to: call a machine learning model to obtain reply data based on the locked preview screen and the audio data collected by the electronic device in real time; and provide a reply result to the user based on the reply data.

[0122] As an exemplary embodiment, the video conversation unit 300 may be configured to: call a machine learning model to obtain reply data based on the real-time preview screen and the audio data collected in real time by the electronic device using the machine learning model; and provide a reply result to the user based on the reply data.

[0123] As an exemplary embodiment, the interface display unit 100 may also be configured to: in response to a user's drawing operation on the locked preview screen, superimpose and display drawing traces formed by the drawing operation on the locked preview screen; wherein, the video conversation unit 300 may be configured to: conduct the real-time video conversation based on the locked preview screen superimposed with the drawing traces.

[0124] As an exemplary embodiment, the interface display unit 100 may also be configured to: in response to a user's drawing operation on the real-time preview screen, overlay and display drawing traces formed by the drawing operation on the real-time preview screen; wherein the video conversation unit 300 may be configured to: conduct the real-time video conversation based on the real-time preview screen overlaid with the drawing traces.

[0125] It should be understood that the specific processing performed by the real-time video conversation device of the electronic device according to the exemplary embodiment of the present disclosure has been referred to. Figures 1 to 12 The details have been described in detail and will not be repeated here.

[0126] In addition, it should be understood that the various units in the real-time video conversation device of the electronic device according to the exemplary embodiment of the present disclosure can be implemented as hardware components and / or software components. Those skilled in the art can implement the various units, for example, using a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC) according to the processing performed by the various units defined.

[0127] According to the computer-readable storage medium of the exemplary embodiment of the present disclosure, a computer program is stored which, when executed by a processor, causes the processor to perform the real-time video conversation method of the electronic device as described in the above exemplary embodiment. The computer-readable storage medium can be any data storage device that can store data read out by a computer system. Examples of computer-readable storage media may include: read-only memory, random access memory, read-only optical disk, magnetic tape, floppy disk, optical data storage device, and carrier wave (such as data transmission through the Internet via a wired or wireless transmission path).

[0128] An electronic device according to an exemplary embodiment of the present disclosure includes: a processor (not shown) and a memory (not shown), wherein the memory stores a computer program, and when the computer program is executed by the processor, the real-time video conversation method of the electronic device as described in the above exemplary embodiment is implemented.

[0129] According to an exemplary embodiment of the present disclosure, a computer program product may also be provided. Instructions in the computer program product may be executed by at least one processor to implement the real-time video conversation method of the electronic device as described in the above exemplary embodiment.

[0130] Although some exemplary embodiments of the present disclosure have been shown and described, it will be appreciated by those skilled in the art that modifications may be made to these embodiments without departing from the principles and spirit of the present disclosure, the scope of which is defined by the claims and their equivalents.

Claims

1. A real-time video conversation method for an electronic device, characterized in that: include: In response to a preset operation, displaying an interface for a user to have a real-time video conversation with an artificial intelligence assistant of the electronic device; In response to a freeze frame function of the real-time video conversation being turned on, determining a locked preview frame of the real-time video conversation; Displaying the locked preview screen in a specific manner in the interface; The real-time video conversation is performed based on the locked preview screen.

2. The real-time video conversation method according to claim 1, characterized in that: Also includes: In response to the freeze function of the real-time video conversation being turned off, displaying a real-time preview of the real-time video conversation in the interface; The real-time video conversation is performed based on the real-time preview image.

3. The real-time video conversation method according to claim 1, characterized in that: In response to the freeze function of the real-time video conversation being turned on, the step of determining the locked preview picture of the real-time video conversation includes: In response to the freeze function of the real-time video conversation being turned on, taking a screenshot of the current preview screen of the real-time video conversation, and using the obtained screenshot as the locked preview screen; Alternatively, in response to the freeze function of the real-time video conversation being turned on, controlling the camera of the electronic device to capture an image, and using the captured image as the locked preview image; Alternatively, in response to the freeze frame function of the real-time video conversation being turned on, candidate preview frames are provided to the user, and the preview frame selected by the user from the candidate preview frames is used as the locked preview frame, wherein the candidate preview frames include the current preview frame and a predetermined number of historical preview frames adjacent to the current preview frame.

4. The real-time video conversation method according to claim 1, characterized in that: The step of displaying the locked preview screen in a specific manner in the interface includes: In the interface, the locked preview screen is used to replace the real-time preview screen of the real-time video conversation for display; Alternatively, in the interface, the locked preview screen is displayed superimposed on the real-time preview screen of the real-time video conversation.

5. The real-time video conversation method according to claim 1, characterized in that: The step of conducting the real-time video conversation based on the locked preview screen includes: Sending the locked preview image to the server at preset intervals; Sending the audio data collected by the electronic device in real time to the server; Acquire, from the server, response data obtained by the server calling the machine learning model based on the locked preview screen and the audio data collected in real time; Based on the reply data, a reply result is provided to the user.

6. The real-time video conversation method according to claim 1, characterized in that: The step of conducting the real-time video conversation based on the locked preview screen includes: Sending the locked preview screen to the server, and request information for requesting the server to store and continue to use the locked preview screen provided this time; Sending the audio data collected by the electronic device in real time to the server; Acquire, from the server, response data obtained by the server calling the machine learning model based on the locked preview screen and the audio data collected in real time; Based on the reply data, a reply result is provided to the user.

7. The real-time video conversation method according to claim 2, characterized in that: The step of conducting the real-time video conversation based on the real-time preview picture includes: Sending the real-time preview image and the audio data collected in real time by the electronic device to a server; Acquire, from the server, response data obtained by the server calling the machine learning model based on the real-time preview image and the real-time collected audio data; Based on the reply data, a reply result is provided to the user.

8. The real-time video conversation method according to claim 1, characterized in that: The step of conducting the real-time video conversation based on the locked preview screen includes: Calling a machine learning model to obtain response data based on the locked preview screen and the audio data collected in real time by the electronic device; Based on the reply data, a reply result is provided to the user.

9. The real-time video conversation method according to claim 2, characterized in that: The step of conducting the real-time video conversation based on the real-time preview picture includes: Calling a machine learning model to obtain response data based on the real-time preview image and the audio data collected in real time by the electronic device using the machine learning model; Based on the reply data, a reply result is provided to the user.

10. The real-time video conversation method according to claim 1, characterized in that: Also includes: In response to a drawing operation performed by a user on the locked preview screen, a drawing trace formed by the drawing operation is superimposed and displayed on the locked preview screen; The step of conducting the real-time video conversation based on the locked preview screen includes: The real-time video conversation is conducted based on the locked preview screen superimposed with the drawing trace.

11. The real-time video conversation method according to claim 2, characterized in that: Also includes: In response to a drawing operation of a user on the real-time preview screen, a drawing trace formed by the drawing operation is superimposed and displayed on the real-time preview screen; Wherein, the step of conducting the real-time video conversation based on the real-time preview screen includes: The real-time video conversation is conducted based on the real-time preview picture superimposed with the painting traces.

12. A real-time video conversation device for an electronic device, characterized in that: include: an interface display unit, configured to display an interface for a real-time video conversation between a user and an artificial intelligence assistant of the electronic device in response to a preset operation; a screen locking unit, configured to determine a locked preview screen of the real-time video conversation in response to a screen freeze function of the real-time video conversation being turned on; a video conversation unit configured to conduct the real-time video conversation based on the locked preview picture in response to the picture freeze function of the real-time video conversation being turned on; The interface display unit is further configured to display the locked preview screen in a specific manner in the interface in response to a freeze function of the real-time video conversation being turned on.

13. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the real-time video conversation method of an electronic device according to any one of claims 1 to 11 is implemented.

14. An electronic device, characterized in that: The electronic device comprises: processor; The memory stores a computer program, and when the computer program is executed by the processor, the real-time video conversation method of the electronic device as described in any one of claims 1 to 11 is implemented.

15. A computer program product comprising computer executable instructions, characterized in that: When the computer executable instructions are executed by a processor, the real-time video conversation method of an electronic device as described in any one of claims 1 to 11 is implemented.