Interaction method, electronic equipment and computer readable storage medium

By using video calls and user behavior detection technology in remote assistance scenes, the location of the subject is displayed with a specific display effect on the assisted device, the problem that the elderly find it difficult to quickly lock the operating position in the shooting screen with complicated information is solved, and the efficiency and accuracy of remote assistance are improved.

CN120017782APending Publication Date: 2025-05-16HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311536846.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In remote assistance scenarios, it is difficult for the elderly to quickly lock the information position of the children's oral in shooting images with complicated information, resulting in poor remote assistance.

Method used

By displaying the subject indicated by the subject with a specific display effect on the second client of the subject, the subject can quickly determine the corresponding operating position of the subject. The specific method includes making a video call with the second client, detecting user behavior, sending video information to display the position of the subject, and setting an animation effect in the shooting screen and/or adding a display mark.

Benefits of technology

It improves the efficiency of remote assistance, enables the assisted to quickly lock the operating position indicated by the helper, and enhances the accuracy and effectiveness of remote assistance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017782A_ABST
    Figure CN120017782A_ABST
Patent Text Reader

Abstract

The invention provides an interaction method, electronic equipment and a computer readable storage medium. According to the method, the user behavior of the assistant in the remote assistance scene is analyzed, the target position of the assistant for indicating the assisted person to operate in the shot picture is determined, and the target position is shared to the client of the assisted person, so that the client of the assisted person marks the target position in the shot picture. Through the method, the assisted person can quickly lock the target position indicated by the assistant to operate, and the remote assistance efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to an interaction method, an electronic device and a computer-readable storage medium. Background Art

[0002] In the application scenario of remote assistance, for example, the children's electronic devices can receive the captured images captured by the elderly's electronic devices through remote communication. Then, the children can assist the elderly in using the hospital's self-service system through oral guidance based on the captured images to complete registration, payment, recharge and other services.

[0003] Specifically, Figure 1 Schematic diagram of a shooting screen is shown. Figure 1 The elderly's mobile phone 10 and the children's mobile phones (not shown) can run a smooth connection TM ,WeChat TM When assistance is needed, the elderly can aim the camera of the mobile phone 10 at the hospital self-service terminal 30, so that the shooting picture 100a is displayed in the video call interface 101. At this time, the shooting picture 100a can be shared with the mobile phones of the children. Figure 1 An enlarged image of the shot screen 100a is shown on the right. As can be seen from the enlarged image, the hospital self-service terminal 30 in the shot screen 100a displays a large amount of text information, such as department name, expert name, title, time period, fee, etc., and some of the text information is similar or identical.

[0004] In the case where the information in such a shot is complex and highly similar, if the children only verbally instruct the elderly to find a specific piece of information in the shot, for example, verbally instructing the elderly to find "Chinese Medicine Internal Medicine", the elderly may not be able to quickly locate the location of "Chinese Medicine Internal Medicine" in the large amount of information in the shot 100a. This will lead to poor remote assistance results. Summary of the invention

[0005] The present application provides an interactive method, an electronic device and a computer-readable storage medium. In a remote assistance scenario, by displaying a photographed object indicated by an assistant with a specific display effect on the assisted person's second client, the assisted person can quickly determine the operation position corresponding to the photographed object, thereby improving the efficiency of remote assistance.

[0006] In a first aspect, the present application provides an interaction method, applied to a first client, the method comprising: conducting a video call with a second client; displaying a first video interface, the first video interface comprising a first window displaying a shot captured by the second client; detecting a first user behavior, wherein the first user behavior is related to a first subject captured by the second client in the first window; and sending first video information to the second client, wherein the first video information is related to a position of the first subject.

[0007] For example, the first client may be the mobile phone 20 of the child described below, or an instant messaging application running in the mobile phone 20. The second client may be the mobile phone 10 of the elderly, or an instant messaging application running in the mobile phone 10. The first video interface may be a video interface of the instant messaging application described below, for example, the video interface 200, the video interface 400, the video interface 500, and the video interface 600 described below.

[0008] The first window in the first video interface may be a window in the mobile phone 20 capable of displaying the captured image of the mobile phone 10, for example, it may be a window 202 in the mobile phone 20 capable of displaying the captured image 202a of the mobile phone 10, a window 401 displaying the captured image 401a of the mobile phone 10, a window 501 displaying the captured image 501a of the mobile phone 10, or a window 601 displaying the captured image 601a of the mobile phone 10.

[0009] The first photographed object may be an object that the child instructs the elderly to operate in the photographed image of the mobile phone 10. For example, it may be "Chinese medicine internal medicine" that the child instructs the elderly to operate in the photographed image of the mobile phone 10.

[0010] Through the above method, the object that the children instruct the elderly to operate can be displayed on the elderly's mobile phone 10, so that the elderly can directly determine the location where the children instruct them to operate, effectively improving the efficiency of remote assistance.

[0011] In a possible implementation of the first aspect above, the first video information includes first position information of the first photographed object in a first photographed picture displayed in the first window.

[0012] Here, the first video information may be the first position information. The first position information may be the position coordinates of the first photographed object in the first photographed screen. The first photographed screen may be the following photographed screen 202a, photographed screen 401a, photographed screen 501a, photographed screen 601a displayed by the mobile phone 20.

[0013] In a possible implementation of the first aspect, the first video information includes a second shot picture, wherein the second shot picture includes the first shot object with the first display effect. The sending of the first video information to the second client includes: setting an animation effect and / or adding a display mark to the first shot object according to the first position information to obtain the second shot picture; and sending the second shot picture to the second client.

[0014] Here, the first video information may be a second shot picture in which the first shot object is displayed with a first display effect according to the first position information. The second shot picture may be, for example, a shot picture 501a marked with a shot object as described below.

[0015] In a possible implementation of the first aspect, the first display effect includes using an animation effect and / or a display mark to indicate an effect of the first photographed object.

[0016] Here, the animation effect may be a dynamic effect such as flashing, rotating, etc. as described below.

[0017] In a possible implementation of the first aspect above, the display mark includes a display box mark and / or a highlight mark.

[0018] Here, the display box mark can be a combination of various line types and annotation shapes such as straight line type, dotted line type, circle, ellipse, rectangle, triangle, etc. in the following text, and the highlight mark can be a highlight display method achieved by superimposing transparent layers of other colors in the following text.

[0019] In a possible implementation of the first aspect above, the first user behavior includes a first gaze behavior of the user, and the first position information of the first photographed object in the first captured picture displayed in the first window is determined in the following manner: collecting a first facial image corresponding to the user's first gaze behavior; determining the user's first gaze position on the first captured picture based on the first facial image; and determining the first gaze position as the first position information.

[0020] Here, the first facial image may be a facial image of the child looking at the photographed picture collected by the mobile phone 20 after receiving the photographed picture sent by the mobile phone 10. The first gaze position may be the gaze position coordinates (x3, y3) of the child on the photographed picture.

[0021] In a possible implementation of the first aspect above, determining a first gaze position of a user on a first captured screen based on a first facial image includes: determining a gaze vector of the user based on the first facial image and a gaze estimation model, wherein the gaze vector is used to indicate a gaze direction corresponding to the first gaze behavior; determining a plurality of second feature points through a key point detection algorithm based on the first facial image and a rigid body model; determining a first feature point through a PnP algorithm based on the plurality of second feature points; determining the first gaze position based on the first feature point and the gaze vector; wherein the second feature points include the coordinates of the left and right inner corners of the eye, the left and right outer corners of the eye, and the left and right nose wings of the user, and the first feature points include the coordinates of the center of the eyebrows of the user.

[0022] Here, the sight estimation model may be a convolutional neural network model described below. The sight vector may be a sight vector (x0, y0, z0) described below. The rigid body model may be a three-dimensional rigid body model of a human head described below. The second feature point may be the three-dimensional rigid body coordinates of the facial feature points of the child described below. The first feature point may be the eyebrow center coordinates (x1, y1, z1) described below.

[0023] In a possible implementation of the first aspect, determining the first gaze position based on the first feature point and the sight line vector includes: determining a second gaze position (x2, y2) of the user on the first client based on the first feature point and the sight line vector:

[0024]

[0025]

[0026] The sight vector is (x0, y0, z0), and the coordinates of the first feature point are (x1, y1, z1); based on the position and size of the first captured image in the first client, the second gaze position is mapped to the first gaze position.

[0027] Here, the second gaze position may be the gaze point coordinates (x2, y2) of the child on the mobile phone 20 described below.

[0028] In a possible implementation of the first aspect above, sending first video information to a second client includes: obtaining second position information, wherein the second position information is determined in the following manner: capturing a second facial image corresponding to a second gaze behavior of a user; determining a third gaze position of the user on a third captured image based on the second facial image; determining the third gaze position as the second position information; wherein the second gaze behavior includes historical gaze behaviors before the first gaze behavior; and corresponding to a distance between the first position information and the second position information being greater than a distance threshold, sending the first video information to the second client.

[0029] Here, the second position information may be the coordinates of the historical gaze position determined last time in the following text, and the first position information may be the coordinates of the current gaze position in the following text. The method for determining the second position information is the same as the method for determining the first position information, and will not be repeated here. In the above manner, the mobile phone 20 can smooth the gaze position, reducing the influence of factors such as picture jitter and / or human eye line jitter on the gaze position.

[0030] In a possible implementation of the first aspect above, the first user behavior also includes oral behavior, wherein the first position information of the first photographed object in the first captured picture displayed in the first window is determined by: collecting voice information corresponding to the user's oral behavior; determining the user's oral position of the first captured picture based on the voice information; and determining the first position information based on the oral position and the first gaze position.

[0031] Here, the first position information determined based on the spoken position and the first gaze position may be the target position described below. Determining the target position in combination with the spoken behavior effectively improves the accuracy of the target position, thereby helping to further improve the efficiency of remote assistance.

[0032] In a possible implementation of the first aspect above, determining the user's oral position of the first captured screen based on voice information includes: establishing a keyword library corresponding to the first captured screen based on a target detection model, wherein the keyword library includes target detection results of all captured objects in the first captured screen; obtaining a semantic analysis result of the voice information through a semantic analysis model; matching the semantic analysis result with the keyword library to obtain target detection results of one or more second captured objects, wherein the second captured objects include the first captured object; and using one or more position coordinates of the one or more second captured objects in the first captured screen as the oral position.

[0033] Here, the target detection result may be all text information in the shooting picture 202a, including keywords such as "department name", "gastroenterology", "Chinese medicine internal medicine", "endocrinology", and "nutrition consultation". The second shooting object may be a shooting object that includes the first shooting object. For example, the second shooting object may be a shooting object whose multiple text information corresponding to display areas 802b, 802c, 802d, and 802e in the shooting picture 802a is "afternoon".

[0034] In a possible implementation of the first aspect, determining the first position information based on the oral position and the first gaze position includes: taking, among one or more position coordinates corresponding to the oral position, position coordinates that are the same as the first gaze position as the first position information.

[0035] Here, the same position coordinates of the spoken position and the first gaze position can be used as the aforementioned target position. For example, the spoken position can be the display area 802b, display area 802c, display area 802d, and display area 802e corresponding to "afternoon" in the captured image 802a described below. The first gaze position can be the display area 802f of the captured image 802a described below. Furthermore, the display area 802c in the captured image 802a where the spoken position and the gaze position overlap can be used as the target position.

[0036] In a possible implementation of the first aspect, sending the first video information to the second client includes: sending the first video information to the second client through the server.

[0037] Here, the server may be an application server hereinafter.

[0038] In a possible implementation of the first aspect above, the first client and the second client include at least one of an instant messaging application, a conference application, a live broadcast application, a teaching application, a game application, and a video application.

[0039] In a second aspect, the present application provides an interaction method, applied to a second client, the method comprising: conducting a video call with a first client; displaying a second video interface, the second video interface comprising a second window displaying a shot screen of the second client, wherein the second window displays a first subject photographed by the second client; receiving first video information, wherein the first video information is related to a position of the first subject displayed in the second window; and displaying the first subject in the second window with a first display effect based on the first video information.

[0040] Here, the second video interface may be the video interface 203 or the video interface 603 mentioned below. The second window may be the window 205 capable of displaying the shooting picture 205a of the mobile phone 10 or the window 605 capable of displaying the shooting picture 605a of the mobile phone 10. In addition, the specific contents of the first client, the second client, the first shooting object, the first video information, and the first display effect have been introduced above and will not be repeated here.

[0041] In a possible implementation of the second aspect above, the first video information includes first position information of the first photographed object in a first photographed picture displayed in the second window.

[0042] Here, the specific content of the first location information has been introduced in the previous article and will not be repeated here.

[0043] In a possible implementation of the second aspect above, displaying the first photographed object with the first display effect in the second window based on the first video information includes: based on the first position information, displaying the first photographed object with the first display effect in the first photographed picture displayed in the second window.

[0044] In a possible implementation of the second aspect, the first video information includes a second shot picture, wherein the second shot picture includes a first shot object having a first display effect.

[0045] Here, the specific content of the second shooting picture has been introduced in the previous article and will not be repeated here.

[0046] In a possible implementation of the second aspect, displaying the first captured object in the second window with the first display effect based on the first video information includes: displaying the second captured picture in the second window.

[0047] In a possible implementation of the second aspect, the first display effect includes using an animation effect and / or a display mark to indicate an effect of the first photographed object.

[0048] In a possible implementation of the second aspect, the display mark includes a display box mark and / or a highlight mark.

[0049] In a possible implementation of the second aspect, receiving the first video information includes: receiving the first video information sent by the second client from the server.

[0050] Here, the specific content of the server has been introduced in the previous article and will not be repeated here.

[0051] In a possible implementation of the second aspect above, the first client and the second client include at least one of an instant messaging application, a conference application, a live broadcast application, a teaching application, a game application, and an audio and video application.

[0052] In the above manner, the second client can display the photographed object instructed by the user of the first client based on the display effect, so that the user of the second client can quickly determine the operation position, thereby improving the efficiency of remote assistance.

[0053] In a third aspect, the present application provides a first electronic device, comprising: one or more processors; one or more memories; one or more memories storing one or more programs, wherein when one or more programs are executed by one or more processors, the first electronic device executes the interaction method executed by the first client provided by the above-mentioned first aspect and various possible implementations of the above-mentioned first aspect.

[0054] In a fourth aspect, the present application provides a second electronic device, comprising: one or more processors; one or more memories; one or more memories storing one or more programs, wherein when one or more programs are executed by one or more processors, the second electronic device executes the interaction method executed by the second client provided by the above-mentioned second aspect and various possible implementations of the above-mentioned second aspect.

[0055] In a fifth aspect, the present application provides a computer-readable medium having instructions stored thereon, which, when executed on a computer, causes the computer to execute the interaction method provided by the first aspect and various possible implementations of the first aspect, the second aspect and various possible implementations of the second aspect.

[0056] In a sixth aspect, an embodiment of the present application further provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute the interaction method provided by the first aspect and various possible implementations of the first aspect, the second aspect and various possible implementations of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 Shown is a schematic diagram of a shooting picture provided by an embodiment of the present application;

[0058] Figure 2a Shown is a schematic diagram of a remote assistance interaction provided by an embodiment of the present application;

[0059] Figure 2b Shown is a schematic diagram of determining a target position provided by an embodiment of the present application;

[0060] Figure 2c Shown is another interactive schematic diagram of remote assistance provided by an embodiment of the present application;

[0061] Figure 3 The figure is a flow chart of an interactive method based on the gaze behavior of the assisted person provided in an embodiment of the present application;

[0062] Figure 4a Shown is a schematic diagram of a shooting picture provided by an embodiment of the present application;

[0063] Figure 4b Shown is a schematic diagram of another shooting picture provided by an embodiment of the present application;

[0064] Figure 4c The figure shows a schematic diagram of a change of a shooting picture provided by an embodiment of the present application;

[0065] Figure 4d The figure shows a schematic diagram of a display interface of an electronic device during a video call provided by an embodiment of the present application;

[0066] Figure 4e Shown is a schematic diagram of a display interface of another electronic device during a video call provided by an embodiment of the present application;

[0067] Figure 4f The figure shows a schematic diagram of a display interface for hiding private information provided by an embodiment of the present application;

[0068] Figure 4g Shown is a schematic diagram of another display interface for hiding private information provided in an embodiment of the present application;

[0069] Figure 5a Shown is a schematic diagram of determining a gaze position provided by an embodiment of the present application;

[0070] Figure 5b Shown is a schematic diagram of an electronic device displaying a gaze position provided by an embodiment of the present application;

[0071] Figure 5c Shown is a schematic diagram of another electronic device displaying a gaze position provided by an embodiment of the present application;

[0072] Figure 5d Shown is a schematic diagram of an electronic device correcting gaze position provided by an embodiment of the present application;

[0073] Figure 6a Shown is a schematic diagram of an interaction effect provided by an embodiment of the present application;

[0074] Figure 6b Shown is another interactive effect schematic diagram provided by an embodiment of the present application;

[0075] Figure 7 The figure is a flow chart of an interactive method based on the gaze behavior and oral behavior of the assisted person provided in an embodiment of the present application;

[0076] Figure 8 Shown is another schematic diagram of determining a target position provided by an embodiment of the present application;

[0077] Fig. 9 Shown is a schematic structural diagram of an electronic device provided in an embodiment of the present application.

[0078] Fig.10a The figure shows a software structure block diagram of an electronic device provided in an embodiment of the present application;

[0079] Fig.10b The figure shows a software structure block diagram of another electronic device provided in an embodiment of the present application;

[0080] Fig.11a The figure shows a schematic diagram of the software system structure of an electronic device provided in an embodiment of the present application;

[0081] Fig.11b FIG. 1 is a schematic diagram of a software system structure of another electronic device provided in an embodiment of the present application;

[0082] Fig.11c The figure shows a schematic diagram of module interaction of an electronic device provided in an embodiment of the present application;

[0083] Fig.11d Shown is a schematic diagram of module interaction of another electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0084] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0085] It can be understood that the electronic devices mentioned in this application may include but are not limited to mobile phones, tablet computers, desktops, laptops, handheld computers, netbooks, as well as augmented reality (AR) and virtual reality (VR) devices, smart TVs, smart watches and other wearable devices, servers, mobile email devices, car devices, portable game consoles, portable music players, reader devices, televisions embedded or coupled with one or more processors, or other electronic devices capable of accessing the Internet.

[0086] It is understood that the instant messaging applications mentioned in this application may include but are not limited to Changlian TM ,WeChat TM ,QQ TM , DingTalk TM The client mentioned in this application can be the aforementioned electronic device capable of running the instant messaging application, or it can be the aforementioned instant messaging application as well as conference application, live broadcast application, teaching application, game application, video application and other application programs with instant messaging function.

[0087] As mentioned above, in a shooting picture with complex information, for example, Figure 1 In the picture 100a shown, which contains a lot of similar or identical information, it is difficult for the elderly (assisted person) to quickly locate the location of the information spoken by the children (assistants). Therefore, the assistance effect that can be achieved by remote assistance based only on oral guidance is not good.

[0088] To solve this problem, the present application provides an interactive method in a remote assistance scenario, which determines the target position indicated by the assistant in the shooting picture by the assistant through analyzing the user behavior of the assistant, and shares the target position with the client of the assistant, so that the client of the assistant can mark the target position in the shooting picture. This method can enable the assistant to quickly lock the target position indicated by the assistant, thereby improving the efficiency of remote assistance.

[0089] The client may include electronic devices such as mobile phones and tablet computers, or may include a mobile phone or tablet computer installed on the electronic device. TM ,WeChat TM ,QQ TM , DingTalk TM There is no restriction on the client side of instant messaging applications such as .

[0090] The user behavior of the above-mentioned assistant may include the assistant's gaze behavior, verbal behavior, etc. For example, when the assistant is looking at the captured image transmitted by the assisted electronic device based on remote communication, the assistant's facial information can be captured by the assistant's electronic device. Furthermore, the assistant's gaze estimation can be performed based on the facial information to obtain the assistant's gaze position on the captured image. Therefore, the assistant's gaze position can be synchronized to the assisted client as the aforementioned target position. For example, the target position can be marked in the captured image displayed by the assisted electronic device.

[0091] Specifically, Figure 2a The interactive schematic diagram of realizing remote assistance based on analyzing the user behavior of the assistant is shown.

[0092] See also Figure 2a , the child's mobile phone 20 can display the video interface 200 of the instant messaging application, and the video interface 200 includes a window 201 and a window 202. Among them, the window 201 can display the shooting picture of the mobile phone 20, and the window 202 can display the shooting picture 202a of the mobile phone 10. The elderly's mobile phone 10 can display the Changlian TM The video interface 203 may display the same content as the video interface 200. For example, the window 204 in the video interface 203 may also display the shooting picture of the mobile phone 20, and the window 205 in the video interface 203 may display the shooting picture 205a of the mobile phone 10.

[0093] The mobile phone 20 can capture the child's facial information, and can determine that the child's gaze position is "Chinese Medicine Internal Medicine" in the shooting screen 202a based on the facial information. Then, the gaze position is shared as a target position to the mobile phone 10, so that the mobile phone 10 marks "Chinese Medicine Internal Medicine" in the shooting screen 205a of the video interface 203 based on the target position. For example, see Figure 2a The enlarged view of the shooting picture 205a shown on the right side shows that the mobile phone 10 can mark the dotted circle in the display area 205b where the "Chinese Medicine Internal Medicine" of the hospital self-service terminal 30 in the shooting picture 205a is located.

[0094] It can be understood that the marking method of the present application can be that after the mobile phone 20 determines the aforementioned target position, the target position is marked in the shooting picture 202a based on the coordinates of the target position in the shooting picture 202a. Furthermore, the mobile phone 20 sends the shooting picture 202a with the target position marked to the mobile phone 10, so that the mobile phone 10 displays the shooting picture 205a with the target position marked. At this time, the shooting picture 205a can be the aforementioned shooting picture 202a. Alternatively, the marking method of the present application can also be that the mobile phone 20 sends the coordinates of the target position in the shooting picture 202a to the mobile phone 10, and after the mobile phone 10 receives the coordinates, the target position is marked in the display area 205b where the "Chinese Medicine Internal Medicine" of the hospital self-service terminal 30 in the shooting picture 205a is located. Here, no restrictive description is given to the specific marking method.

[0095] It is understandable that when the elderly's electronic device collects and shoots pictures, the pictures may shake due to the unstable state of the elderly's handheld device. Therefore, when the shaking pictures are transmitted to the children's client, the children may not be able to keep their sights stable at the desired position. Furthermore, the target position determined by the children's client may drift or be too large. Alternatively, the shaking of the children's own sight may also cause the target position determined by the children's client to drift or be too large.

[0096] In this case, in order to solve the impact of screen jitter and / or line of sight jitter on the target position, specifically, after the child's mobile phone 20 determines the current target position based on the child's gaze behavior, it can be determined whether to update the target position based on the historical target position determined last time. For example, a distance threshold can be set. If the distance value between the current target position and the historical target position is less than the distance threshold, it can be considered that the change in the current target position may be caused by factors such as jitter in the shooting picture, and the target position that the child intends to instruct the elderly to operate has not changed. At this time, the target position may not be updated, and the historical target position continues to be shared with the mobile phone 10. On the contrary, if the distance value between the current target position and the historical target position is greater than the distance threshold, it can be considered that the target position that the child intends to instruct the elderly to operate has changed. At this time, the current target position can be shared with the mobile phone 10. By smoothing the target position in the above manner, the impact of factors such as screen jitter or line of sight jitter on the determination of the target position can be reduced.

[0097] In other embodiments, when the assistant orally instructs the assisted person based on the shot screen, the assistant's voice information can be semantically analyzed based on the assistant's client, and the shot screen can be image understood to establish a keyword library corresponding to the shot screen. Then, the result of the semantic analysis is matched with the keyword library to determine the target position indicated by the assistant in the shot screen. Then, the determined target position can be synchronized to the assisted person's client, so that the assisted person's client can mark the target position in the displayed shot screen.

[0098] Specifically, continue to refer to the above Figure 2a , the mobile phone 20 can perform semantic analysis on the voice information of the children and determine that the voice information includes "Chinese medicine internal medicine". In addition, the mobile phone 20 can perform image understanding on the shooting screen 202a of the video interface 200 and establish a keyword library of the shooting screen 202a, which may include all the text information in the shooting screen 202a. Then, the result of the semantic analysis is matched with the keyword library, and the matching result may be "Chinese medicine internal medicine". Therefore, the position of "Chinese medicine internal medicine" in the shooting screen 202a can be shared with the mobile phone 10 as the target position, so that the mobile phone 10 marks the display area 205b where "Chinese medicine internal medicine" is located in the shooting screen 205a of the video interface 203 based on the target position. In this way, the mobile phone 10 can also mark the target position in the display area 205b where "Chinese medicine internal medicine" is located based on the marking method described above. The specific marking method can be found in the detailed description above, which will not be repeated here.

[0099] In other embodiments, the target position indicated by the assistant can also be determined based on the assistant's gaze behavior and oral behavior. This method can effectively reduce the adverse effect of image jitter on the accuracy of the target position for the aforementioned scene with image jitter.

[0100] As an example, Figure 2b A schematic diagram of determining a target position based on analyzing the gaze behavior and oral behavior of an assistant is shown. The schematic diagram includes an enlarged view of the shooting picture 202a, a schematic diagram of the gaze position for line of sight estimation of the shooting picture 202a, and a schematic diagram of the target position determined based on line of sight estimation and semantic analysis.

[0101] Figure 2b Taking the example that the mobile phone 20 first determines the gaze position and then determines the target position based on the gaze position and the oral position, the specific process of determining the target position is shown. It can be understood that the mobile phone 20 can also first determine the oral position, and then determine the target position based on the oral position and the gaze position; or the mobile phone 20 can also determine the gaze position and the oral position at the same time, and then determine the overlapping part of the two positions as the target position. Here, this application does not make a restrictive description on the order of determining the gaze position and the oral position, and the schemes for jointly determining the target position based on the gaze position and the oral position are all within the protection scope of this application.

[0102] See also Figure 2b , the mobile phone 20 displays the shooting picture of the mobile phone 20 in the window 201 of the video interface 200, and displays the shooting picture 202a collected by the mobile phone 10 in the window 202. The children want to instruct the elderly to click on the "Chinese Medicine Internal Medicine" position of the self-service system based on the shooting picture 202a, but due to the shaking of the picture, the children's line of sight drifts in the larger area containing "Chinese Medicine Internal Medicine" in the shooting picture 202a of the video interface 200. As a result, when the children's mobile phone 20 estimates the line of sight, it may use the display area 202b containing "Chinese Medicine Internal Medicine", "Gastroenterology", and "Endocrinology" as the gaze position. If the gaze position is sent as the target position, the accuracy of remote assistance will be reduced.

[0103] Therefore, the mobile phone 20 can further combine the children's voice information and the semantic analysis results in the display area 202b to determine a more accurate target position. For example, the mobile phone 20 can determine the keywords contained in the display area 202b in the aforementioned keyword library, and match them with the semantic analysis results of the children's voice information to use the display area 202c of "Chinese Medicine Internal Medicine" in the shooting screen 202a as the target position.

[0104] It can be understood that the aforementioned method of the mobile phone 20 determining the target position in the display area 202b in combination with the semantic analysis result is only an example, and the mobile phone 20 can also determine the target position in the shooting screen 202a in combination with the semantic analysis result. In addition, the mobile phone 20 can also first match the semantic analysis result with the keyword library, and then determine the target position based on the matching result combined with the gaze position. The order of semantic analysis and line of sight estimation is not limited here.

[0105] Figure 2c According to an embodiment of the present application, an interactive schematic diagram of realizing remote assistance based on analyzing the gaze behavior and oral behavior of the assistant is shown.

[0106] See also Figure 2c When the mobile phone 20 determines the aforementioned display area 202c as the target location, the target location can be shared with the mobile phone 10, so that the mobile phone 10 marks the display area 205b where "Chinese Medicine Internal Medicine" is located in the shooting screen 205a of the video interface 201 based on the target location. In this way, the mobile phone 10 can also mark the target location in the display area 205b where "Chinese Medicine Internal Medicine" is located based on the marking method as described above. The specific marking method can be found in the detailed description above, which will not be repeated here.

[0107] Based on the above Figure 1 The remote assistance scenario shown, combined with the embodiments and drawings, introduces in detail the interaction solution provided by this application example.

[0108] Specifically, the specific implementation process of the interaction method provided by the present application based on the gaze behavior of the assisted person is first introduced below in conjunction with an embodiment.

[0109] Example 1

[0110] This embodiment will explain in detail the specific implementation process of the interaction method based on the gaze behavior of the assisted person in conjunction with the accompanying drawings.

[0111] Figure 3 According to an embodiment of the present application, a flow chart of an interaction method implemented based on the gaze behavior of the assisted person is shown.

[0112] Understandably, Figure 3 The execution subject of the interactive process shown can be the client of the assisted person and the client of the assisting person. The following takes the client of the assisted person as the mobile phone 10 of the elderly and the client of the assisting person as the mobile phone 20 of the children as an example to describe the interactive solution provided in the embodiment of the present application in detail.

[0113] Specifically, Figure 3 As shown, the method may include the following steps:

[0114] 301: Mobile phone 10 and mobile phone 20 run an instant messaging application.

[0115] In one example, an instant messaging application is used as a link TM For example, both the mobile phone 10 and the mobile phone 20 are running Changlian TM Application, to establish a remote communication connection based on the following step 302. It can be understood that Changlian TM ,WeChat TM Instant messaging applications such as WhatsApp are essentially client programs installed on electronic devices such as mobile phones. In other embodiments, electronic devices such as mobile phones can also run web client versions of instant messaging applications, which are not limited here.

[0116] 302: Mobile phone 10 establishes a remote communication connection with mobile phone 20.

[0117] In an exemplary manner, the mobile phone 10 and the mobile phone 20 may establish a remote communication connection by starting a video call. TM The application initiates a video call request to the child's mobile phone 20, and the Changlian running on the mobile phone 20 TM After the application responds to the request, the mobile phone 10 establishes a remote communication connection with the mobile phone 20, and can transmit video call data based on the application server, etc. The video data can be the captured images and / or voice data collected by the mobile phone 10 and / or the mobile phone 20.

[0118] In other embodiments, the children can also use the mobile phone 20 to run the Changlian TM The application initiates a video call request to the elderly person’s mobile phone 10. TM After the application responds to the request, the mobile phone 10 establishes a remote communication connection with the mobile phone 20. Furthermore, the mobile phone 10 and the mobile phone 20 can implement the following steps 306 to 308 based on the remote communication connection.

[0119] It is understood that the above application server can be a mobile phone 10 or mobile phone 20 running Changlian TM In other embodiments, remote communication and transmission of video call data between mobile phone 10 and mobile phone 20 may also be achieved through other servers, such as general servers, database servers or file servers, etc., which are not limited here.

[0120] Here, it can be understood that the interactive method provided by this application can only be implemented when one party initiates a video call request and the other party responds to the request. Therefore, the assistant cannot obtain the information of the assisted electronic device under the condition that the assisted person does not agree to remote assistance, which effectively protects the privacy of the assisted person.

[0121] 303: The mobile phone 10 captures the captured image.

[0122] In an example manner, the elderly can use the mobile phone 10 to shoot content that requires remote assistance from their children. Here, the content of remote assistance can be equipment that requires guidance on operation, etc. In addition, based on the foregoing content, it can be seen that when the elderly use the mobile phone 10 to shoot the hospital self-service terminal 30, the captured picture may shake because they cannot hold the device continuously and stably. Therefore, the mobile phone 10 can capture the captured picture in combination with the camera anti-shake technology to reduce the shaking problem of the captured picture. Here, the camera anti-shake technology includes but is not limited to optical image stabilization (OIS) technology and electronic image stabilization (EIS) technology.

[0123] As an example, Figure 4a A schematic diagram of a shooting screen is shown.

[0124] See also Figure 4a In the shooting picture shown, the elderly may need their children to guide how to use the hospital's self-service system to register for a specialist clinic. In this scenario, the elderly can focus the rear camera of the mobile phone 10 on the hospital self-service terminal 30 of the hospital, so that the shooting picture of the mobile phone 10 displays the "specialist clinic registration" page of the hospital self-service terminal 30.

[0125] Figure 4b A schematic diagram showing another shooting picture is shown.

[0126] See also Figure 4b In the shooting screen shown, the elderly may need their children to guide how to use the hospital's self-service system to make an examination appointment. In this scenario, the elderly can focus the rear camera of the mobile phone 10 on the hospital's self-service terminal 30, so that the shooting screen of the mobile phone 10 displays multiple function selection buttons of the hospital self-service terminal 30. The aforementioned function selection buttons include but are not limited to "booking number", "registration", "recharge", "charge", "query", "booking", and "examination appointment".

[0127] 304: The mobile phone 10 sends the captured image to the mobile phone 20 based on the remote communication connection.

[0128] Exemplarily, the mobile phone 10 can synchronize the shooting picture of the hospital self-service terminal 30 to the mobile phone 20 by means of a video call. Here, if the shooting picture collected by the mobile phone 10 based on the camera anti-shake technology in the aforementioned step 303 still has a shaking phenomenon, for example, the shooting picture has edge distortion and / or edge blur, etc., at this time, the mobile phone 10 can correct the shooting picture. For example, the mobile phone 10 can perform edge trimming processing on the shooting picture, and use the central area of ​​the picture that is not affected by the shaking after trimming as the transmitted shooting picture. Then, the mobile phone 20 displays the shooting picture after the edge trimming processing. It can be understood that this method can effectively improve the display effect of the shooting picture by the mobile phone 20. In addition, the mobile phone 20 can determine the gaze position of the children based on the shooting picture after the edge trimming processing in the following step 305. In this way, the mobile phone 10 can send the shooting picture with the shaking effect reduced to the mobile phone 20, thereby reducing the influence of the picture shaking on the target position, which is conducive to improving the accuracy of the subsequent determination of the gaze position.

[0129] Specifically, the mobile phone 10 may perform edge trimming on the captured image before sending it to the mobile phone 20; or the mobile phone 20 may perform edge trimming on the captured image after receiving the captured image from the mobile phone 10. The following takes the correction of the captured image by the mobile phone 10 as an example to specifically describe the edge trimming operation performed by it. It can be understood that if the mobile phone 20 performs edge trimming on the captured image, the edge trimming operation performed by the mobile phone 20 is substantially the same as the edge trimming operation performed by the mobile phone 10, and will not be described in detail here.

[0130] For example, Figure 4c A schematic diagram showing a variation of edge trimming processing performed by a mobile phone 10 is shown.

[0131] See also Figure 4c , because the elderly cannot hold the device continuously and stably, the captured image 401a captured by the mobile phone 10 may shake. For example, the hospital self-service terminal 30 captured in the captured image 401a may have edge distortion and / or edge blur. In this case, the mobile phone 10 can trim the edge of the captured image 401a based on a preset ratio to obtain a trimmed captured image 401b. It can be understood that since the captured image 401b has eliminated the influence of image shaking, after the mobile phone 10 sends the captured image 401b to the mobile phone 20, the mobile phone 20 can determine a more accurate gaze position based on the following step 305 without being affected by image shaking.

[0132] In one example approach, Figure 4d A schematic diagram of a display interface of a mobile phone 20 during a video call is shown.

[0133] See also Figure 4dAfter establishing a communication connection with the mobile phone 10, the mobile phone 20 displays a video interface 400. The video interface 400 may include a window 401 capable of displaying a shooting picture 401a of the mobile phone 10, and a window 402 capable of displaying a shooting picture of the mobile phone 20. Figure 4d Take the example of window 401 being tiled on the video interface 400 and window 402 being suspended on the upper layer of window 401 in a small window mode. It can be understood that this window display mode is only an example, and in other display modes, for example, window 402 can be tiled on the video interface 400, and window 401 can be suspended on the upper layer of window 402 in a small window mode. This application does not make a restrictive description of the display mode of window 401 and window 402.

[0134] Here, the captured image 401a captured by the mobile phone 10 may be jittery, so the mobile phone 20 may display the captured image 401b after edge trimming.

[0135] In another example, Figure 4e Another schematic diagram of the display interface of the mobile phone 20 during a video call is shown.

[0136] See also Figure 4e , the captured image 401a captured by the mobile phone 10 may not have jitter, so the mobile phone 20 can directly display the captured image 401a without edge trimming. Figure 4e The window display method in Figure 4d The same, no further elaboration is given here.

[0137] In addition, if the captured image captured by the mobile phone 10 contains privacy information, the mobile phone 10 and / or the mobile phone 20 may perform privacy processing on the captured image captured by the mobile phone 10 .

[0138] Specifically, Figure 4f A display schematic diagram of a mobile phone 20 is shown.

[0139] See also Figure 4f , the captured image 401c captured by the mobile phone 10 includes the hospital self-service terminal 30 displaying the "Fee Inquiry" interface, and the "Fee Inquiry" interface includes the "Hospitalization ID" and "Password" which are private information. Therefore, in an example manner, the private information in the captured image 401c can be hidden by the mobile phone 10. Specifically, after the mobile phone 10 detects that there is private information in its captured image 401c, the "Hospitalization ID" and "Password" in the captured image 401c can be blurred. For example, the mobile phone 10 can superimpose an opaque layer on the display area of ​​the private information. Then, the mobile phone 10 sends the blurred captured image to the mobile phone 20, so that the mobile phone 20 displays the blurred captured image 401c in the window 401.

[0140] In another example, the mobile phone 20 may hide the private information in the shooting picture 401c. Specifically, after detecting that there is hidden information in the shooting picture 401c sent by the mobile phone 10, the mobile phone 20 may blur the private information in the shooting picture 401c. For example, the mobile phone 20 may overlay an opaque layer on the display area of ​​the private information. Then, the mobile phone 20 may display the blurred shooting picture 401a in the window 401.

[0141] Specifically, Figure 4g Another display schematic diagram of a mobile phone 20 is shown.

[0142] See also Figure 4g After detecting that there is hidden information in the shooting picture 401c sent by the mobile phone 10, the mobile phone 20 can display the window 401 in the video interface 400 in a black screen.

[0143] The foregoing Figure 4f and Figure 4g The two fuzzy processing methods shown can effectively prevent the leakage of the privacy information of the elderly. Here, the privacy information may include but is not limited to fingerprint information, palm print information, account information, password information, QR code information, bar code information, etc. It can be understood that the aforementioned Figure 4f and Figure 4g The two blurring processing methods shown are only examples. This application does not make any restrictive descriptions on the specific blurring processing methods. All methods that can make private information invisible in the mobile phone 20 are within the protection scope of this application.

[0144] 305: The mobile phone 20 determines the user's gaze position on the captured image based on the user's gaze behavior.

[0145] In an example manner, after receiving the captured image sent by the mobile phone 10 based on the above step 304, the mobile phone 20 can collect the facial image of the child when looking at the captured image. For example, the mobile phone 20 can turn on the front camera to capture the facial image of the child, and estimate the line of sight of the child based on the facial image. As an example, the mobile phone 20 can perform line of sight estimation based on a pre-trained convolutional neural networks (CNN) model. For example, the facial image of the child can be used as data input into the CNN model to generate a line of sight vector (x0, y0, z0) representing the line of sight direction. Here, the CNN model can be, for example, a residual network (ResNet) model, etc., and this application does not restrict the specific model type for line of sight estimation.

[0146] Furthermore, the mobile phone 20 can obtain the coordinates of the child's facial feature points based on the collected child's facial image. Here, the facial feature points may include the inner corner of the eye, the outer corner of the eye, the nose wings on both sides, etc. This application does not restrict the specific positions of the facial feature points. The feature points that can be used to determine the starting point of the child's line of sight belong to the facial feature points described in this application.

[0147] Specifically, the mobile phone 20 can use the facial image of the child and the three-dimensional rigid body model of the human head as input data, and obtain the three-dimensional rigid body coordinates of the facial feature points of the child through a keypoint alignment algorithm. The aforementioned three-dimensional rigid body model and three-dimensional rigid body coordinates are both located in the camera coordinate system of the mobile phone 20, which is a three-dimensional rectangular coordinate system established with the focus center of the front camera of the mobile phone 20 as the origin and the optical axis as the Z axis.

[0148] Furthermore, the coordinates of the child's eyebrow center (x1, y1, z1) can be determined based on the three-dimensional rigid coordinates of the child's facial feature points by solving the 3D point to 2D point (perspective-n-point, PnP) algorithm. Here, the eyebrow center coordinates can be used as the starting point of the aforementioned sight vector. Furthermore, the mobile phone 20 can calculate the coordinates of the child's gaze point on the mobile phone 20 (x2, y2) based on the sight vector (x0, y0, z0) and the eyebrow center coordinates (x1, y1, z1):

[0149]

[0150]

[0151] Based on the above, the mobile phone 20 can determine the gaze position coordinates of the children on the shooting screen of the mobile phone 10 according to the above gaze point coordinates (x2, y2). For example, if the mobile phone 20 displays the video interface in full screen on the display screen, and the mobile phone 20 displays the shooting screen of the mobile phone 10 in full screen on the video interface, it can be understood that the gaze point coordinates (x2, y2) at this time are the gaze position coordinates (x3, y3) of the children on the shooting screen of the mobile phone 10. Alternatively, if the shooting screen of the mobile phone 10 is not displayed in full screen on the display screen of the mobile phone 20, the gaze point coordinates can be mapped to the shooting screen based on the position and size of the window where the shooting screen is located, and the gaze position coordinates (x3, y3) of the children on the shooting screen are obtained.

[0152] As an example, Figure 5a A schematic diagram of determining a gaze position is shown.

[0153] See also Figure 5a, the mobile phone 20 displays the shooting picture 501a shot by the mobile phone 10, and the shooting picture 501a displays the "specialist outpatient registration" interface of the hospital self-service terminal 30. Because the children want to instruct the elderly to click on "Chinese medicine internal medicine" in the "specialist outpatient registration" interface, they continue to look at the location of "Chinese medicine internal medicine" in the shooting picture 501a. Based on the above method, the mobile phone 20 estimates and determines the gaze position coordinates in the "Chinese medicine internal medicine" area in the shooting picture 501a, and then, based on the gaze position coordinates, the mobile phone 20 determines the gaze position 501b of the children on the shooting picture 501a.

[0154] In addition, the mobile phone 20 can also smooth the determined gaze position coordinates to avoid the gaze position determined based on the gaze position coordinates from being offset and / or jittering due to factors such as picture jitter and / or line of sight jitter. Specifically, the mobile phone 20 can pre-set a distance threshold. After the mobile phone 20 determines the current gaze position coordinates based on the above process, it can calculate the distance between the current gaze position coordinates and the historical gaze position coordinates determined last time. If the distance value is less than the distance threshold, the mobile phone 20 may consider that the change in the current gaze position coordinates may be caused by factors such as jitter in the shooting picture, and the actual gaze position of the children has not changed. Therefore, the gaze position may not be updated, and the historical gaze position coordinates continue to be shared to the mobile phone 10 based on the following step 305. On the contrary, if the distance value is greater than the distance threshold, the mobile phone 20 may consider that the actual gaze position of the children has changed. Therefore, the mobile phone 20 may share the current gaze position coordinates to the mobile phone 10 based on the following step 306. In the above manner, the mobile phone 20 smoothes the gaze position, reducing the influence of factors such as picture jitter and / or line of sight jitter on the gaze position.

[0155] 306 : Mobile phone 20 sends the gaze position to mobile phone 10 .

[0156] In an example manner, the mobile phone 20 may send the gaze position coordinates corresponding to the gaze position to the mobile phone 10, so that the mobile phone 10 displays the photographed object corresponding to the gaze position in the photographing picture based on the following step 307.

[0157] In another example, the mobile phone 20 may also display the gaze position in the captured image, and then send the captured image showing the gaze position to the mobile phone 10. Specifically, after the mobile phone 20 determines the gaze position in the captured image of the mobile phone 10, the captured object corresponding to the gaze position may be displayed in the captured image. For example, the mobile phone 20 may use the display area centered on the gaze position in the captured image as the captured object corresponding to the gaze position, and mark the display area in the captured image in a preset manner.

[0158] As an example, Figure 5bA schematic diagram showing a mobile phone 20 displaying a gaze position is shown.

[0159] See also Figure 5b , the video interface 500 of the mobile phone 20 includes a window 501 for displaying the shooting picture 501a of the mobile phone 10, and a window 502 capable of displaying the shooting picture of the mobile phone 20. After the mobile phone 20 determines the gaze position 501b in the shooting picture 501a of the mobile phone 10, the display area 501c centered on the gaze position 501b in the shooting picture 501a can be used as the shooting object corresponding to the gaze position 501b. Furthermore, the mobile phone 20 can mark the display area 501c in the shooting picture 501a in the form of a dotted ellipse. It can be understood that the size of the display area centered on the gaze position can be preset based on the actual scene, and this application does not make a restrictive description on the size of the display area.

[0160] As an example, Figure 5c Another schematic diagram of a mobile phone 20 displaying a gaze position is shown.

[0161] See also Figure 5c , when the mobile phone 20 determines the gaze position 501b in the shooting picture 501a of the mobile phone 10, the display area corresponding to the gaze position 501b can also be used as its corresponding shooting object. Furthermore, the mobile phone 20 can directly mark the gaze position 501b in the shooting picture 501a in the form of a solid line dot. It can be understood that the way to mark the gaze position 501b and / or the display area 501c can be a combination of multiple line types and marking shapes such as a straight line type, a dotted line type, a circle, an ellipse, a rectangle, a triangle, etc., and this application does not make a restrictive description on the specific marking method.

[0162] Here, since the children can see the photographed object marked by the mobile phone 20 in the photographed screen 501a, the children can judge whether the marking result of the mobile phone 20 is correct. If the marking result is wrong, the marking result can be corrected, and the corrected gaze position can be sent to the mobile phone 10.

[0163] As an example, Figure 5d A schematic diagram of a mobile phone 20 correcting gaze position is shown.

[0164] See also Figure 5d , if the mobile phone 20 selects the display area 501d of "Nutrition Consultation Department" as the shooting object when the child is looking at "Chinese Medicine Internal Medicine", the woman can drag the dotted ellipse annotation box to the display area 501c of "Chinese Medicine Internal Medicine" by dragging with one finger to correct the annotation result. It can be understood that the child can also move the annotation box by single-clicking, double-clicking the correct gaze position, dragging the annotation box to the correct gaze position with two fingers, etc. This application does not make a restrictive description of the specific user operation of moving the annotation box.

[0165] Here, it can be understood that the children's mobile phone 20 can only send the children's gaze position to the elderly's mobile phone 10, but cannot remotely control the mobile phone 10, thus ensuring the privacy and security of the elderly.

[0166] 307: The mobile phone 10 displays the photographed object corresponding to the gaze position in the photographing screen.

[0167] In an example, after the mobile phone 10 receives the gaze position coordinates corresponding to the gaze position sent by the mobile phone 20, the mobile phone 10 can display the photographed object corresponding to the gaze position in the photographed screen of the mobile phone 10 according to the gaze position coordinates. For example, the mobile phone 10 can use the display area centered on the gaze position coordinates in the photographed screen as the photographed object corresponding to the gaze position, and mark the display area in the photographed screen in a preset manner.

[0168] Specifically, after receiving the gaze position coordinates (x3, y3) sent by the mobile phone 20, the mobile phone 10 may map the gaze position coordinates (x3, y3) sent by the mobile phone 20 into the window according to the size of the window displaying the shooting picture, and obtain the mapped gaze position coordinates (x4, y4):

[0169]

[0170]

[0171] Wherein, w1 is the window width of the mobile phone 20 displaying the captured image;

[0172] h1 is the height of the window of the mobile phone 20 displaying the captured image;

[0173] w2 is the width of the window on which the mobile phone 10 displays the captured image;

[0174] h2 is the height of the window of the mobile phone 10 displaying the captured image.

[0175] As an example, Figure 6a A schematic diagram of an interaction effect is shown.

[0176] See also Figure 6a The video interface 600 of the mobile phone 20 includes a window 601 that displays a shooting screen 601a of the mobile phone 10, and a window 602 that can display the shooting screen of the mobile phone 20. The video interface 603 of the mobile phone 10 includes a window 605 that displays a shooting screen 605a of the mobile phone 10, and a window 604 that can display the shooting screen of the mobile phone 20.

[0177] Specifically, the elderly's mobile phone 10 makes a video call with the children's mobile phone 20. When the mobile phone 20 sends the gaze position coordinates (x3, y3) corresponding to the gaze position to the mobile phone 10, the mobile phone 10 can map the gaze position coordinates in the window 605 to obtain the mapped gaze position coordinates (x4, y4). Furthermore, the mobile phone 10 can use the display area 605b centered on the gaze position coordinates (x4, y4) in the shooting screen 605a as the shooting object corresponding to the gaze position. Furthermore, the display area 605b can be marked in the shooting screen 605a in the form of a dotted ellipse. This allows the elderly to directly see the location where the children indicate the operation in the window 605, achieve rapid positioning in the remote assistance scene, and improve the efficiency of remote assistance. In addition, the mobile phone 10 can also directly use the display area corresponding to the gaze position coordinates as its corresponding shooting object, and mark the display area of ​​the gaze position coordinates in the shooting screen 605a. The display effect is the same as the aforementioned. Figure 5c The display effects shown are substantially the same and are not shown here.

[0178] Here, the way in which the mobile phone 10 marks the gaze position coordinates and / or the display area 605b can be a combination of a plurality of line types and annotation frame shapes such as a straight line type, a dotted line type, a circle, an ellipse, a rectangle, a triangle, etc., and the present application does not make a restrictive description of the specific annotation method. For example, the mobile phone 10 can also highlight the display area 605b, for example, the screen color of the display area 605b can be changed, or a transparent layer of other colors can be superimposed on the display area 605b. The mobile phone 10 can also set a dynamic effect for the display area 605b, for example, the annotation frame of the display area 605b can be set to flash, rotate and other dynamic effects. The mobile phone 10 can also display the display area 605b in an enlarged manner. It can be understood that the display area 605b can be distinguished from other display areas, so that the display area 605b is eye-catching and easy to distinguish. All methods are within the scope of protection of this application.

[0179] It is understood that during the remote interaction between the mobile phone 10 and the mobile phone 20, the mobile phone 10 and / or the mobile phone 20 may close the shooting screen of the mobile phone 20 and only display the shooting screen 601a and / or the shooting screen 605a of the mobile phone 10. For example, the mobile phone 10 may close the window 604 and only display the window 605 on the video interface 603, and / or the mobile phone 20 may close the window 602 and only display the window 601 on the video interface 600. Alternatively, the mobile phone 10 and / or the mobile phone 20 may only hide the shooting screen of the mobile phone 20, for example, the mobile phone 10 may display the default screen in the window 604 / or the mobile phone 20 may display the default screen in the window 602. Here, the default screen may be a headshot picture of the children, a solid color background picture, etc. The specific content of the default screen is not limited here.

[0180] For another example, Figure 6bAnother schematic diagram of the interaction effect is shown.

[0181] Specifically, the elderly's mobile phone 10 makes a video call with the children's mobile phone 20. When the mobile phone 20 sends the shooting picture 601a with the shooting object marked to the mobile phone 10, the mobile phone 10 can directly display the shooting picture 601a in the window 605. Then, the elderly can directly see the location where the children indicate the operation in the window 605, which improves the efficiency of remote assistance. Figure 6b The captured image 605a displayed by the mobile phone 10 is the captured image 601a captured and displayed by the mobile phone 20 .

[0182] In addition, the display area corresponding to the gaze position coordinates can also be directly used as the photographed object corresponding to the gaze position, so as to mark the display area of ​​the gaze position coordinates in the photographing screen 601a and the photographing screen 605a. The display effect of the mobile phone 10 and the mobile phone 20 is the same as that of the aforementioned Figure 5c The display effects shown are substantially the same and are not shown here.

[0183] The above embodiment 1 can display the photographed object that the child intends to instruct the elderly to operate on the elderly's mobile phone 10 based on the child's gaze behavior. The specific implementation process of implementing the interactive method provided by this application based on the child's gaze behavior and oral behavior will be further described in conjunction with embodiment 2.

[0184] Example 2

[0185] This embodiment will explain in detail the specific implementation process of the interaction method based on the gaze behavior and oral behavior of the assisted person in conjunction with the accompanying drawings.

[0186] Figure 7 According to the embodiment of the present application, a flow chart of an interactive method based on the gaze behavior and oral behavior of the assisted person is shown. It can be understood that Figure 7 The execution subject of the interactive process shown can be the client of the assisted person and the client of the assisting person. The following takes the client of the assisted person as the mobile phone 10 of the elderly and the client of the assisting person as the mobile phone 20 of the children as an example to describe the interactive solution provided in the embodiment of the present application in detail.

[0187] Specifically, Figure 7 As shown, the method may include the following steps:

[0188] 701: Mobile phone 10 and mobile phone 20 run an instant messaging application.

[0189] Here, the content of the instant messaging application running on the mobile phone 10 and the mobile phone 20 may be specifically described in detail in the aforementioned step 301, and will not be elaborated here.

[0190] 702: Mobile phone 10 establishes a remote communication connection with mobile phone 20.

[0191] Here, the details of establishing a remote communication connection between the mobile phone 10 and the mobile phone 20 can be specifically referred to the detailed description in the aforementioned step 302, which will not be repeated here.

[0192] 703: The mobile phone 10 captures the captured image.

[0193] Here, the mobile phone 10 collects the content of the captured image. For details, please refer to the detailed description in the aforementioned step 303, which will not be repeated here.

[0194] 704 : The mobile phone 10 sends the captured image to the mobile phone 20 based on the remote communication connection.

[0195] Here, the mobile phone 10 sends the content of the captured image to the mobile phone 20 based on the remote communication connection. For details, please refer to the detailed description in the aforementioned step 304, which will not be repeated here.

[0196] 705: The mobile phone 20 determines the user's target position on the captured image based on the user's gaze behavior and oral behavior.

[0197] For example, after receiving the captured image sent by the mobile phone 10 based on the above step 703, the mobile phone 20 can determine the child's gaze position on the captured image based on the child's gaze behavior on the captured image. Here, the mobile phone 20 determines the user's gaze position on the captured image based on the user's gaze behavior. For details, please refer to the detailed description in the above step 304, which will not be repeated here.

[0198] Furthermore, the mobile phone 20 can capture the user's oral behavior through the microphone during the video call. For example, the mobile phone 20 can collect the voice information of the children. Then, the mobile phone 20 can input the voice information into the semantic understanding model, and determine the children's oral position of the shot screen in combination with the semantic analysis results and the image understanding results of the shot screen. Therefore, the mobile phone 20 can determine the target position that the children instruct the elderly to operate based on the gaze position and the oral position.

[0199] Specifically, the aforementioned Figure 2b A schematic diagram of a process for determining a target position is shown.

[0200] In an exemplary manner, the mobile phone 20 can generate a corresponding keyword library by performing image understanding on the shot image 202a based on the target detection model. For example, the keyword library corresponding to the shot image 202a can include all text information in the shot image 202a, including keywords such as "department name", "gastroenterology", "TCM internal medicine", "endocrinology", and "nutrition consultation". Furthermore, the mobile phone 20 can match the semantic analysis results of the child's voice information with the keywords in the aforementioned gaze position.

[0201] For example, the gaze position determined by the mobile phone 20 based on the line of sight analysis is the display area 202b. Based on the optical character recognition (OCR) of the display area 202b and the matching with the aforementioned keyword library, it can be determined that the keywords included therein are "gastroenterology", "TCM Internal Medicine", and "endocrinology". The voice information collected by the mobile phone 20 may be "click on the location of TCM Internal Medicine", and its semantic analysis result may be "click", "TCM Internal Medicine", and "location". Furthermore, the mobile phone 20 matches the keywords in the display area 202b with the aforementioned semantic analysis results to obtain the matching result "TCM Internal Medicine". Furthermore, the mobile phone 20 can use the display area 202c corresponding to "TCM Internal Medicine" as the target location for children to instruct the elderly to operate.

[0202] It can be understood that, here, the mobile phone 20 may use the display area 202b containing multiple department names as the gaze position due to the shaking of the shooting screen 202a and / or the drifting of the children's line of sight when looking at the shooting screen 202a. This gaze position cannot directly lock the position where the children instruct the elderly to operate. However, in combination with the oral position, the display area 202c containing only the keyword "Chinese Medicine Internal Medicine" can be used as the target position, which effectively corrects the gaze position, improves the accuracy of the target position, and thus helps to further improve the efficiency of remote assistance.

[0203] Figure 8 According to an embodiment of the present application, another schematic diagram of a process for determining a target position is shown.

[0204] See also Figure 8 , the mobile phone 20 receives the shooting picture 802a captured by the mobile phone 10. In the process of determining the target position of the shooting picture 802a, the mobile phone 20 may collect the voice information of the child based on the microphone, which may be "select the afternoon time", and the result of semantic analysis of the voice information may be "select", "afternoon", "time". The semantic analysis result is matched with the keyword library of the shooting picture 802a, and the keyword "afternoon" can be obtained. Furthermore, the mobile phone 20 can use the display area 802b, display area 802c, display area 802d, and display area 802e corresponding to the keyword "afternoon" as the oral position based on the text recognition of the shooting picture 802a.

[0205] In addition, the mobile phone 20 can use the display area 802f as the gaze position based on the line of sight analysis of the captured image 802a. Furthermore, the display area 802c where the spoken position and the gaze position overlap can be used as the target position.

[0206] 706 : Mobile phone 20 sends the target location to mobile phone 10 .

[0207] Here, the mobile phone 20 sends the content of the target location to the mobile phone 10. For details, please refer to the detailed description in the aforementioned step 306, which will not be repeated here.

[0208] 707: The mobile phone 10 displays the photographed object corresponding to the target position in the photographing screen.

[0209] Here, the mobile phone 10 displays the content of the photographed object corresponding to the target position in the photographed screen. For details, please refer to the detailed description in the aforementioned step 307, which will not be repeated here.

[0210] It can be understood that the embodiment of the present application provides an interactive solution based on the above steps 701 to 706 that combines gaze behavior and oral behavior to jointly determine the target position, which is conducive to eliminating interference such as shooting picture jitter, user line of sight drift, etc., and improving the accuracy of the corresponding determined target position.

[0211] Fig. 9 According to an embodiment of the present application, a structural schematic diagram of an electronic device 100 is shown.

[0212] Here, the electronic device 100 may be, for example, the aforementioned mobile phone 10 or mobile phone 20, etc., which is not limited here.

[0213] like Fig. 9 As shown, the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0214] It is to be understood that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figure, or combine some components, or separate some components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.

[0215] The processor 110 may include one or more processing units, for example, the processor 110 may include an application processor (AP), a modem processor, a graphics processor (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0216] In some embodiments, the processor 110 may include one or more interfaces. The interface may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0217] The MIPI interface can be used to connect the processor 110 with peripheral devices such as the display screen 194 and the camera 193. The MIPI interface includes a camera serial interface (CSI), a display serial interface (DSI), etc. In some embodiments, the processor 110 and the camera 193 communicate via the CSI interface to implement the shooting function of the electronic device 100. The processor 110 and the display screen 194 communicate via the DSI interface to implement the display function of the electronic device 100.

[0218] The wireless communication function of the electronic device 100 can be realized by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor, etc. The electronic device 100 can realize the aforementioned remote video communication based on the wireless communication function.

[0219] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve the utilization of antennas. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0220] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc., applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In some embodiments, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.

[0221] The modem processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be sent into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After the low-frequency baseband signal is processed by the baseband processor, it is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to a speaker 170A, a receiver 170B, etc.), or displays an image or video through a display screen 194. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.

[0222] The wireless communication module 160 can provide wireless communication solutions including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc., which are applied to the electronic device 100. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, modulates the frequency of the electromagnetic wave signal and performs filtering, and sends the processed signal to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, modulate the frequency of it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0223] In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technology. The above wireless communication technology may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The above-mentioned GNSS may include a global positioning system (GPS), a global navigation satellite system (GLONASS), a Beidou navigation satellite system (BDS), a quasi-zenith satellite system (QZSS) and / or a satellite based augmentation system (SBAS).

[0224] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), Mini-LED, Micro-LED, Micro-OLED, quantum dot light-emitting diodes (QLED), etc. In some embodiments, the electronic device 100 may include 1 or N display screens 194, where N is a positive integer greater than 1. Exemplarily, the electronic device 100 can display the aforementioned video interface and the captured image through the display screen 194.

[0225] The electronic device 100 can realize the shooting function through ISP, camera 193, video codec, GPU, display screen 194 and application processor.

[0226] ISP is used to process the data fed back by camera 193. For example, when taking a photo, the shutter is opened, and the light is transmitted to the camera photosensitive element through the lens. The light signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to ISP for processing and converts it into an image visible to the naked eye. ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. ISP can also optimize the exposure, color temperature and other parameters of the shooting scene. In some embodiments, ISP can be set in camera 193.

[0227] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1. Exemplarily, the electronic device 100 may capture the aforementioned captured image based on the camera 193.

[0228] NPU is a neural network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission mode between neurons in the human brain, it can quickly process input information and can also continuously self-learn. Through NPU, applications such as intelligent cognition of electronic device 100 can be realized, such as image recognition, face recognition, voice recognition, text understanding, etc.

[0229] The electronic device 100 can implement audio functions such as music playing and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0230] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be arranged in the processor 110, or some functional modules of the audio module 170 can be arranged in the processor 110.

[0231] The speaker 170A, also called a "speaker", is used to convert an audio electrical signal into a sound signal. The electronic device 100 can listen to music or listen to a hands-free call through the speaker 170A.

[0232] The receiver 170B, also called a "earpiece", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or voice message, the voice can be received by placing the receiver 170B close to the human ear.

[0233] Microphone 170C, also called "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to microphone 170C to input the sound signal into microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the sound source, realize directional recording function, etc.

[0234] The pressure sensor 180A is used to sense the pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be set on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. A capacitive pressure sensor can be a parallel plate including at least two conductive materials. When a force acts on the pressure sensor 180A, the capacitance between the electrodes changes. The electronic device 100 determines the intensity of the pressure based on the change in capacitance. When a touch operation acts on the display screen 194, the electronic device 100 detects the intensity of the touch operation based on the pressure sensor 180A. The electronic device 100 can also calculate the touch position based on the detection signal of the pressure sensor 180A. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities can correspond to different operation instructions.

[0235] The gyro sensor 180B can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., x, y, and z axes) can be determined by the gyro sensor 180B. The gyro sensor 180B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyro sensor 180B detects the angle of the electronic device 100 shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the electronic device 100 through reverse movement to achieve anti-shake. The gyro sensor 180B can also be used for navigation and somatosensory game scenes.

[0236] The touch sensor 180K is also called a "touch control device". The touch sensor 180K can be set on the display screen 194. The touch sensor 180K and the display screen 194 form a touch screen, also called a "touch control screen". The touch sensor 180K is used to detect touch operations acting on or near it. The touch sensor can pass the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor 180K can also be set on the surface of the electronic device 100, which is different from the position of the display screen 194.

[0237] Fig.10a According to an embodiment of the present application, a software structure block diagram of a mobile phone 10 is shown.

[0238] The software system of the mobile phone 10 can adopt a layered architecture, an event-driven architecture, a micro-kernel architecture, a micro-service architecture, or a cloud architecture. TM Taking the system as an example, the software structure of the mobile phone 10 is exemplified.

[0239] The layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, Android TM The system is divided into four layers, from top to bottom: application layer, application framework layer, Android TM Runtime (Android TM runtime) and system libraries, as well as the kernel layer.

[0240] like Fig.10a As shown, the application layer may include a series of application packages. The application package may include applications such as instant messaging applications.

[0241] The instant messaging application of the mobile phone 10 may include an image acquisition module 1001, a communication module 1002, and a display module 1003, which are used to execute the above-mentioned embodiment 1. Figure 3 Or in Example 2 Figure 7 For the sake of narrative coherence, the specific functions of the image acquisition module 1001, the communication module 1002, and the display module 1003 will be described in detail below. Fig.11a and Fig.11c Specific instructions are given in the description.

[0242] As mentioned above, in other embodiments, the instant messaging application installed on the mobile phone 10 may be Changlian. TM ,WeChat TM ,QQ TM And other applications.

[0243] The application framework layer provides an application programming interface (API) and a programming framework for the applications in the application layer. The application framework layer includes some predefined functions.

[0244] like Fig.10a As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.

[0245] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.

[0246] Content providers are used to store and retrieve data and make it accessible to applications. This data can include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.

[0247] The view system includes visual controls, such as controls for displaying text, controls for displaying images, etc. The view system can be used to build applications. A display interface can be composed of one or more views. For example, a display interface including a text notification icon can include a view for displaying text and a view for displaying images.

[0248] The phone manager is used to provide communication functions of the mobile phone 10, such as management of call status (including connecting, hanging up, etc.).

[0249] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.

[0250] The notification manager enables applications to display notification information in the status bar. It can be used to convey notification-type messages and can disappear automatically after a short stay without user interaction. For example, the notification manager is used to notify download completion, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as notifications of applications running in the background, or a notification that appears on the screen in the form of a dialog window. For example, it can prompt text information in the status bar, emit a prompt sound, vibrate, or flash an indicator light.

[0251] Android TM Runtime includes core libraries and virtual machines. Android TM The runtime is responsible for scheduling and management of the Android system.

[0252] The core library consists of two parts: one part is the function that needs to be called by the Java language, and the other part is the Android core library.

[0253] The application layer and the application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object life cycle management, stack management, thread management, security and exception management, and garbage collection.

[0254] The system library may include multiple functional modules, such as surface manager, media libraries, 3D graphics processing library (such as openGL ES), 2D graphics engine (such as SGL), etc.

[0255] The surface manager is used to manage the display subsystem and provide the fusion of 2D and 3D layers for multiple applications.

[0256] The media library supports playback and recording of a variety of commonly used audio and video formats, as well as static image files, etc. The media library can support a variety of audio and video encoding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.

[0257] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0258] A 2D graphics engine is a drawing engine for 2D drawings.

[0259] The kernel layer is the layer between hardware and software. The kernel layer contains at least display driver, camera driver, audio driver, and sensor driver.

[0260] Fig.10b According to an embodiment of the present application, a software structure block diagram of a mobile phone 20 is shown.

[0261] The software structure of the mobile phone 20 can refer to the detailed description of the software structure of the mobile phone 10 mentioned above, which will not be described in detail here.

[0262] Different from the mobile phone 10, the instant messaging application of the mobile phone 20 may include a sight focus recognition module 2001, a communication module 2002, and a display module 2003, which are used to perform the above Figure 3 and / or Figure 7 For the sake of narrative coherence, the specific functions of the sight focus recognition module 2001, the communication module 2002, and the display module 2003 will be described in detail below. Fig.11b and Fig.11d Specific instructions are given in the description.

[0263] Fig.11a A schematic diagram of the software system structure of a mobile phone 10 is shown.

[0264] See also Fig.11a The mobile phone 10 includes an image acquisition module 1001 , a communication module 1002 , and a display module 1003 .

[0265] The image acquisition module 1001 is used to obtain the image captured by the mobile phone 10 .

[0266] The communication module 1002 is used to establish remote communication between the mobile phone 10 and the mobile phone 20, for example, to implement a video call between the mobile phone 10 and the mobile phone 20.

[0267] The display module 1003 is used to display the aforementioned video interface, including displaying the shooting screen of the mobile phone 10 and the shooting screen of the mobile phone 20 .

[0268] Here, the image acquisition module 1001 , the communication module 1002 , and the display module 1003 may be functional modules in an instant messaging application running on the mobile phone 10 .

[0269] Fig.11b A schematic diagram of the software system structure of a mobile phone 20 is shown.

[0270] See also Fig.11b The mobile phone 20 includes a sight focus recognition module 2001 , a communication module 2002 , and a display module 2003 .

[0271] The sight focus recognition module 2001 is used to obtain the picture taken by the mobile phone 20, and determine the gaze position of the child on the picture based on the picture.

[0272] The communication module 2002 is used to establish remote communication between the mobile phone 10 and the mobile phone 20, for example, to implement a video call between the mobile phone 10 and the mobile phone 20.

[0273] The display module 2003 is used to display the aforementioned video interface, including displaying the shooting screen of the mobile phone 10 and the shooting screen of the mobile phone 20 .

[0274] Here, the sight focus recognition module 2001 , the communication module 2002 , and the display module 2003 may be functional modules in an instant messaging application running on the mobile phone 10 .

[0275] Fig.11c A schematic diagram of the calling relationship between hardware and software in the system of a mobile phone 10 is shown.

[0276] See also Fig.11c The physical unit of the mobile phone 10 may include a rear camera, a central processing unit (CPU), and a wireless chip. The rear camera may be the aforementioned camera 193, which is used to shoot a picture including content requiring remote assistance. The CPU may be the aforementioned processor 110. The wireless chip may be a chip used for remote video communication in the aforementioned mobile communication module 150 and / or wireless communication module 160.

[0277] The functional modules of the mobile phone 10 may include the aforementioned image acquisition module 1001 and the aforementioned communication module 1002 .

[0278] Here, the image acquisition module 1001 can be used to call the rear camera and the CPU, and the communication module 1002 can be used to call the CPU and the wireless chip.

[0279] Fig.11d A schematic diagram of the calling relationship between hardware and software in a system of a mobile phone 20 is shown.

[0280] See also Fig.11d The physical unit of the mobile phone 20 may include a front camera, a CPU, and a wireless chip. The front camera may be the aforementioned camera 193, which is used to shoot a picture including content requiring remote assistance. The CPU may be the aforementioned processor 110. The wireless chip may be a chip used for remote video communication in the aforementioned mobile communication module 150 and / or wireless communication module 160.

[0281] The functional modules of the mobile phone 20 may include a sight focus recognition module 2001 and a communication module 2002 .

[0282] Here, the sight focus recognition module 2001 can be used to call the front camera and the CPU, and the communication module 2002 can be used to call the wireless chip and the CPU.

[0283] The embodiments of the present application also provide a computer program product for implementing the interaction methods provided in the above embodiments.

[0284] The various embodiments of the mechanism disclosed in the present application can be implemented in hardware, software, firmware or a combination of these implementation methods. The embodiments of the present application can be implemented as a computer program module or module code executed on a programmable system, and the programmable system includes at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device and at least one output device.

[0285] A computer program module or module code can be applied to input instructions to perform the functions described in this application and generate output information. The output information can be applied to one or more output devices in a known manner. For the purposes of this application, a processing system includes any system having a processor such as, for example, a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0286] Module code can be implemented with high-level modular language or object-oriented programming language to communicate with the processing system. When necessary, module code can also be implemented with assembly language or machine language. In fact, the mechanism described in this application is not limited to the scope of any specific programming language. In either case, the language can be a compiled language or an interpreted language.

[0287] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented as instructions carried or stored on one or more temporary or non-temporary machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors. For example, instructions may be distributed over a network or through other computer-readable media. Therefore, machine-readable media may include any mechanism for storing or transmitting information in a machine (e.g., computer) readable form, including but not limited to floppy disks, optical disks, optical disks, read-only memories (CD-ROMs), magneto-optical disks, read-only memories (ROM), random access memories (RAM), erasable programmable read-only memories (EPROM), electrically erasable programmable read-only memories (EEPROM), magnetic or optical cards, flash memory, or a tangible machine-readable memory for transmitting information (e.g., carrier waves, infrared signals, digital signals, etc.) using the Internet in electrical, optical, acoustic, or other forms of propagation signals. Accordingly, machine-readable media include any type of machine-readable media suitable for storing or transmitting electronic instructions or information in a form readable by a machine (eg, a computer).

[0288] References to "one embodiment" or "an embodiment" in the specification mean that the specific features, structures, or characteristics described in conjunction with the embodiment are included in at least one exemplary implementation or technology disclosed according to the embodiment of the present application. The appearance of the phrase "in one embodiment" in various places in the specification does not necessarily all refer to the same embodiment.

[0289] The disclosure of the embodiment of the present application also relates to an operating device for executing the text. The device can be specially constructed for the required purpose or it can include a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program can be stored in a computer-readable medium, such as, but not limited to any type of disk, including a floppy disk, an optical disk, a CD-ROM, a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an EPROM, an EEPROM, a magnetic or optical card, an application-specific integrated circuit (ASIC) or any type of medium suitable for storing electronic instructions, and each can be coupled to a computer system bus. In addition, the computer mentioned in the specification may include a single processor or may be an architecture involving multiple processors for increased computing power.

[0290] In addition, the language used in this specification has been primarily selected for readability and instructional purposes and may not be selected to describe or limit the disclosed subject matter. Therefore, the present application embodiment disclosure is intended to illustrate rather than limit the scope of the concepts discussed herein.

Claims

1. An interaction method, applied to a first client, characterized in that: The method includes: Conduct a video call with a second client; Displaying a first video interface, wherein the first video interface includes a first window displaying a picture captured by the second client; A first user behavior is detected, wherein the first user behavior is related to a first photographed object photographed by the second client in the first window; Sending first video information to the second client, wherein the first video information is related to the position of the first photographed object.

2. The method according to claim 1, characterized in that The first video information includes first position information of the first shooting object in a first shooting picture displayed in the first window.

3. The method according to claim 2, characterized in that The first video information includes a second shooting picture, wherein the second shooting picture includes the first shooting object with a first display effect.

4. The method according to claim 3, characterized in that The first display effect includes an effect of using an animation effect and / or a display mark to indicate the first photographed object.

5. The method according to claim 4, characterized in that The display mark includes a display frame mark and / or a highlight mark.

6. The method according to claim 4, characterized in that The sending the first video information to the second client includes: According to the first position information, The first photographed object is set with an animation effect and / or the display mark is added to obtain the second photographed picture; Send the second captured image to the second client.

7. The method according to claim 2, characterized in that The first user behavior includes a first gaze behavior of the user, and, The first position information of the first photographed object in the first photographed picture displayed in the first window is determined by: Collecting a first facial image corresponding to the first gaze behavior of the user; determining a first gaze position of the user on the first captured image based on the first facial image; The first gaze position is determined as the first position information.

8. The method according to claim 7, characterized in that The determining, based on the first facial image, a first gaze position of the user on the first captured picture includes: determining a gaze vector of the user based on the first facial image and a gaze estimation model, wherein the gaze vector is used to indicate a gaze direction corresponding to the first gaze behavior; Determine a plurality of second feature points by a key point detection algorithm based on the first facial image and the rigid body model; Determine the first feature point by using a PnP algorithm based on the plurality of second feature points; Determining the first gaze position based on the first feature point and the sight line vector; The second feature points include the coordinates of the left and right inner corners of the eyes, the coordinates of the left and right outer corners of the eyes, and the coordinates of the left and right nose wings of the user. The first feature point includes the coordinates of the center of the user's eyebrows.

9. The method according to claim 8, characterized in that The determining the first gaze position based on the first feature point and the sight line vector includes: Determine a second gaze position (x2, y2) of the user on the first client based on the first feature point and the sight line vector: Wherein, the sight line vector is (x0, y0, z0), and the coordinates of the first feature point are (x1, y1, z1); The second gaze position is mapped to the first gaze position based on the position and size of the first captured image in the first client.

10. The method according to claim 7, characterized in that , the sending of the first video information to the second client includes: Acquire second location information, wherein the second location information is determined in the following manner: collecting a second facial image corresponding to a second gaze behavior of the user; determining a third gaze position of the user on a third captured picture based on the second facial image; Determining the third gaze position as the second position information; Wherein, the second gaze behavior includes historical gaze behaviors before the first gaze behavior; Corresponding to the distance between the first position information and the second position information being greater than a distance threshold, the first video information is sent to the second client.

11. The method according to claim 7, characterized in that The first user behavior also includes a verbal behavior, wherein the first position information of the first photographed object in the first photographed picture displayed in the first window is determined by: Collecting voice information corresponding to the user's oral behavior; Determining the user's spoken position of the first captured image based on the voice information; The first position information is determined based on the spoken position and the first gaze position.

12. The method according to claim 11, characterized in that The determining, based on the voice information, the user's spoken position of the first photographed picture includes: Establishing a keyword library corresponding to the first shooting picture based on the target detection model, wherein the keyword library includes target detection results of all shooting objects in the first shooting picture; Obtaining a semantic analysis result of the speech information through a semantic analysis model; Matching the semantic analysis result with the keyword library to obtain target detection results of one or more second photographed objects, wherein the second photographed objects include the first photographed object; The one or more position coordinates of the one or more second photographed objects in the first photographed picture are used as the spoken position.

13. The method according to claim 12, characterized in that The determining the first position information based on the oral position and the first gaze position includes: Among the one or more position coordinates corresponding to the spoken position, the position coordinate that is the same as the first gaze position is used as the first position information.

14. The method according to any one of claims 1 to 13, characterized in that The sending the first video information to the second client includes: The first video information is sent to the second client through the server.

15. The method according to any one of claims 1 to 14, characterized in that The first client and the second client include at least one of an instant messaging application, a conference application, a live broadcast application, a teaching application, a game application, and a video application.

16. An interaction method, applied to a second client, characterized in that: The method includes: Conducting a video call with the first client; Displaying a second video interface, wherein the second video interface includes a second window displaying a shot image of the second client, wherein the second window displays a first shot object photographed by the second client; receiving first video information, wherein the first video information is related to a position of the first photographed object displayed in the second window; The first photographed object is displayed in the second window with a first display effect based on the first video information.

17. The method according to claim 16, characterized in that The first video information includes first position information of the first shooting object in the first shooting picture displayed in the second window.

18. The method according to claim 17, characterized in that The displaying the first photographed object in the second window with a first display effect based on the first video information includes: Based on the first position information, the first photographing object is displayed in the first photographing picture displayed in the second window with a first display effect.

19. The method according to claim 17, characterized in that The first video information includes a second shooting picture, wherein the second shooting picture includes the first shooting object with a first display effect.

20. The method according to claim 19, characterized in that The displaying the first photographed object in the second window with a first display effect based on the first video information includes: The second shooting picture is displayed in the second window.

21. The method according to claim 18 or 19, characterized in that The first display effect includes an effect of using an animation effect and / or a display mark to indicate the first photographed object.

22. The method according to claim 21, characterized in that The display mark includes a display frame mark and / or a highlight mark.

23. The method according to any one of claims 17 to 22, characterized in that The receiving of the first video information comprises: The first video information sent by the second client is received from the server.

24. The method according to any one of claims 17 to 23, characterized in that The first client and the second client include at least one of an instant messaging application, a conference application, a live broadcast application, a teaching application, a game application, and an audio-visual application.

25. A first electronic device, characterized in that: include: one or more processors; One or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the first electronic device executes any one of the interaction methods of claims 1 to 15 executed by the first client.

26. A second electronic device, characterized in that: include: one or more processors; One or more memories; the one or more memories store one or more programs, and when the one or more programs are executed by the one or more processors, the second electronic device executes any one of claims 16 to 24 of the interactive method executed by the second client.

27. A computer readable medium, characterized in that The readable medium stores instructions, which, when executed on a computer, cause the computer to execute the interactive method according to any one of claims 1 to 24.