Method and device for constructing a knowledge base for the purpose of making cross-functional use of the application functions of a plurality of software items

EP4569410A1Pending Publication Date: 2025-06-18ORANGE SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
EP2023748811
Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-10
Filing Date
2023-08-02
Publication Date
2025-06-18

AI Technical Summary

Technical Problem

The increasing complexity of user interfaces across multiple applications on electronic terminals leads to a lack of homogeneity and user confusion, as each application offers its own unique experience, making it difficult for users to adapt and control multiple applications simultaneously without a unified interface.

Method used

A method and device that construct a knowledge base by analyzing system events, screen captures, and context data to establish a unified interface, allowing an Over The Top (OTT) application to manage multiple applications independently of their APIs, using cursor position, system, and context data to simulate user interactions and provide a unified user experience.

Benefits of technology

Enables users to control multiple applications with a single interface, reducing adaptation time and allowing seamless interaction across different applications, even with updates or new installations, without requiring access to each application's APIs, thus providing a generic and adaptive solution for managing various applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 1.1
    Figure 1.1
Patent Text Reader

Abstract

The invention relates to a method for constructing a knowledge base, implemented by a construction device during use of an electronic terminal, and characterized in that the method comprises: when a system event on said electronic terminal is detected - a first step of obtaining at least one position datum in relation to a cursor, said cursor being associated with at least one pointing peripheral, and at least one digital image of a snapshot of at least part of the output of at least one screen of said terminal; - a second step of obtaining at least one system datum from said electronic terminal; - a third step of obtaining at least one context datum based on analysis of all or part of said digital image; - a step of updating said knowledge base with said at least one position datum, system datum and context datum.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DESCRIPTION

[0002] Title: Method and device for constructing a knowledge base with the aim of using application functions of a plurality of software programs in a cross-functional manner.

[0003] 1. Field of the invention

[0004] The invention relates to the field of electronic terminals capable of executing a plurality of applications. More particularly, the invention relates to techniques allowing, for example, the execution of functions of one or more applications in a transverse manner, that is to say without depending on a given application.

[0005] 2. Prior Art

[0006] Electronic terminals (computers, smartphones, tablets, etc.) can have increasingly larger screens and computing capacities allowing them to run numerous computer applications simultaneously.

[0007] In this context, each application can offer its own unique user experience. One disadvantage is that this can lead to a lack of consistency between the user interfaces of the applications running on the electronic terminal.

[0008] Additionally, some applications evolve regularly to offer new features. This is the case, for example, when applications offer APIs (Application Programming Interface) that allow the development of extension / external modules (plugins) capable of providing new features. A disadvantage is that application user interfaces are increasingly rich / complex and therefore difficult for the user to understand.

[0009] So, when a user uses an application for the first time, it is not uncommon for them to need a more or less long adaptation period before being able to use it correctly.

[0010] There is therefore a need for a simple solution allowing a user and / or a dedicated application to control, with a unified user experience, a plurality of applications executed on an electronic terminal, the solution having to be independent of the applications to be controlled. 3. Presentation of the invention

[0011] The invention improves the state of the art and proposes for this purpose a method for constructing a knowledge base, said method being implemented by a construction device during use of an electronic terminal, and characterized in that the method comprises: when a system event on said electronic terminal is detected

[0012] - a first step of obtaining at least one item of cursor position data, said cursor being associated with at least one pointing device of said electronic terminal, and at least one digital image of a capture of at least one part of the rendering of at least one screen of said electronic terminal;

[0013] - a second step of obtaining at least one system data from said electronic terminal;

[0014] - a third step of obtaining at least one piece of context data based on an analysis of all or part of said digital image;

[0015] - a step of updating said knowledge base with said at least one position, system and context data.

[0016] Thus, the proposed solution is based on a new and inventive approach consisting of building a knowledge base by exploiting both the system events of an electronic terminal triggered by a user but also screen capture and image analysis technologies in order to automatically establish a table of correspondences between each function of each application (computer application) used by the user and the associated system actions (actions executed by the electronic terminal).

[0017] An advantage of the proposed solution is that it is simple to implement since all that is required, in addition to the terminal already available to the user, is a computing machine (possibly the one already present in the terminal).

[0018] Another advantage of the proposed solution is that it allows a generic construction of the knowledge base by freeing itself from access to the APIs of each application executed by the terminal. In other words, because it is not based on the APIs of an application executed by the terminal, the proposed solution allows the creation of a knowledge base independent ("agnostic") of this application, with regard to the way of collecting information related to the use of this application. Therefore, the proposed solution requires few implementation constraints.

[0019] Another advantage of the proposed solution is that a computer application configured to use the content of this knowledge base (OTT application for Over The TOP in English or so-called bypass application) can be used to control transversely, via for example the same human-machine interface, the different computer applications present on the electronic terminal. Indeed, when the user requests the execution of an action / function on his terminal via the OTT application, for example the action of launching an IP (Internet Protocol) telephony application, it consults the knowledge base and then retrieves, depending on the requested action (for example the label), the information necessary to execute the action.The information may include: the position of the icon of an IP telephony application, i.e. the position of the cursor obtained and saved within the knowledge base during a previous use / execution of the IP telephony application (for example when the user selects / clicks on the icon of the IP telephony application with a mouse connected to the electronic terminal). According to a particular embodiment, the position of the cursor saved within the knowledge base corresponds to the position of the icon when the user releases the pressure on a mouse button of his electronic terminal. This embodiment makes it possible, when the user moves the icon of the IP telephony application, to update the knowledge base with the new position of the icon.Indeed, the movement of an icon on a screen of an electronic terminal is for example carried out by a non-released click, that is to say held down, on a mouse button of the electronic terminal combined with a movement of the cursor to the desired position. Once the position is reached, the user releases the pressure exerted on the mouse button so that the icon is placed in the desired position;

[0020] - context data such as the size and position of the user interface of the telephony application to be executed (execution parameters); etc. Once this data has been retrieved, the OTT application can then execute the telephony application, for example, by simulating a click on the icon of the IP telephony application. The OTT application can also provide display parameters to the IP telephony application once it is started.

[0021] In other words, the OTT computing application does not have to be application-specific and can cooperate with the generic knowledge base, which may contain information related to multiple applications (even if in a particular implementation it may also contain information related to a single application).

[0022] Furthermore, even if the application(s) evolve (for example via an update and a change of version), or if the user adds an application to his terminal, the proposed solution continues to function without requiring an update, since it relies on screen extractions (partial or total) and / or system events.

[0023] Furthermore, the knowledge base is enriched over time thanks to user actions carried out at the level of the electronic terminal applications.

[0024] Note that the method may, when obtaining context data and / or system data, detect that the application used by the user corresponds to the OTT application. In this case, the method may not update the knowledge base.

[0025] According to a particular embodiment, the OTT application can autonomously and transversely control the functions of the applications of the electronic terminal according to predetermined computer routines (a sequence of computer instructions). A routine can for example comprise the detection of a particular event obtained from the electronic terminal such as the detection of a user action, the exceeding of a threshold or a duration, etc.

[0026] An electronic terminal is any device capable of at least managing a display device and / or a pointing / input device (personal computer, smartphone, electronic tablet, television, on-board computer of a car, connected objects, etc.). A system event is an event generated by an operating system of an electronic terminal. The system event is, for example, generated upon receipt of a message or following an action by a user at the level of a peripheral of the electronic terminal. A system event can, for example, be triggered following the execution of a computer command (initiated or not by the user).

[0027] System data means data obtained, for example, from the operating system of the electronic terminal.

[0028] According to a particular embodiment of the invention, a method as described above is characterized in that said system event is generated by at least one pointing device associated with said electronic terminal.

[0029] In this embodiment, the method is triggered when the user interacts and performs an action (e.g., a click) at a pointing device associated with the electronic terminal.

[0030] A pointing device is any input device that allows a user to enter position data (coordinates / spatial data), for example via a cursor, and / or action data, for example via a click, at an electronic terminal. A pointing device is, for example, a touchpad, a mouse, a trackball, a trackpoint or a joystick.

[0031] According to a particular embodiment of the invention, a method as described above is characterized in that the capture relates to an active application window.

[0032] In this embodiment, the method captures a portion of the screen that corresponds to the active application window displayed on the screen. An application window is a window linked to an execution of an application by the terminal. An active application window is an application window that is currently being used by the user, i.e., which holds the focus. This embodiment applies in particular in the case where the terminal allows multi-windowing (i.e., can simultaneously display several application windows).

[0033] If the terminal can only display one application window at a time (e.g., a smartphone terminal), the application window corresponds to the active application window. The capture is then performed on the entire terminal screen.

[0034] Note that the terminal may be associated with or include one or more display devices.

[0035] According to a particular embodiment of the invention, a method as described above is characterized in that the obtaining of said at least one piece of context data is carried out via an optical character recognition technique and / or a computer vision technique.

[0036] In this way, the knowledge base can be enriched with two types of information: those extracted from the text and those extracted from the image elements. This covers most, or in some cases all, of the useful data included in the image.

[0037] According to a particular embodiment of the invention, a method as described above is characterized in that the updating step is conditioned by the value of a confidence score associated with said at least one piece of context data.

[0038] In this way, the quality of the information collected and stored in the knowledge base is improved.

[0039] According to a particular embodiment of the invention, a method as described above is characterized in that said at least one system data item comprises at least one computer command executable by the operating system of said electronic terminal.

[0040] In this way, the information collected and stored in the knowledge base includes system commands that can be replayed by an OTT computer application. For example, when the user wishes to program the shutdown of his Windows 10™ computer via the OTT application, the latter obtains the associated command from the knowledge base (shutdown -s -f -t xxx where "xxx" corresponds to the desired delay). Obviously, this assumes that this action has already been carried out beforehand by the user via another application and added to the knowledge base. An advantage of this embodiment is that it is possible to control an application and / or trigger a computer function even when it is not accessible via the human-machine interface provided by the electronic terminal.Indeed, the execution of the command makes it possible to execute the function requested by the user without having to simulate a mouse click on the graphical button or a menu associated with the requested function. The system data may also include the name of the active application (i.e. the application currently being used by the user). The name of the application is for example obtained from the operating system of the electronic terminal via a system command or via a specific API such as the JavaScript Node .js command.

[0041] ",getActiveWindow()" from the "npm" package manager. This is generally the name of the application's executable file, that is, the file containing the computer code that allows the application to be executed by the electronic terminal. Note that the name of the executable can be compared to elements of a list containing the commercial names of the applications. Thus, it is possible, using the name of the executable, to obtain the name of the application whose graphical interface is displayed on a screen of the electronic terminal.

[0042] According to a particular embodiment of the invention, a method as described above is characterized in that said context data belongs to the group comprising at least:

[0043] - a label;

[0044] - a graphics window size;

[0045] - a position within said screen;

[0046] - a text;

[0047] - an image.

[0048] Thus, the proposed solution can take into account the great diversity of data obtained as a result of the analysis of the digital image. It is effective even if the user manipulates a large quantity of applications. Concretely, the context data can include the version of the application, the name of the application / computer function, the description of the function, the position within the image of a graphic element and / or a label associated with the function, and more generally any information associated with the application and / or the computer function used by the user at the electronic terminal. The context data can also include a keyboard shortcut associated with the function used and / or an image symbolizing the function.

[0049] The context data may also include the nature / type of the graphical window displayed by a computer application on a screen of the electronic terminal. Indeed, it is for example possible to distinguish, using computer vision techniques, a videoconferencing window from an instant messaging window displayed by the same application (for example Microsoft Teams). Thus, the nature / type of the window may correspond to the main function rendered by the graphical window (editing an email, videoconferencing, instant messaging, document database, etc.).

[0050] This list of context data types is not exhaustive.

[0051] The invention also relates to a device for constructing a knowledge base implemented during use of an electronic terminal, and characterized in that the device comprises:

[0052] - a first module for obtaining at least one item of cursor position data, said cursor being associated with said at least one pointing device of said electronic terminal, and at least one digital image of a capture of at least one part of the rendering of at least one screen of said electronic terminal;

[0053] - a second module for obtaining at least one system data item from said electronic terminal;

[0054] - a third module for obtaining at least one piece of context data based on an analysis of all or part of said digital image;

[0055] - a module for updating said knowledge base with said at least one position, system and context data.

[0056] The term module can correspond to a software component as well as to a hardware component or a set of hardware and software components, a software component itself corresponding to one or more computer programs or subroutines or more generally to any element of a program capable of implementing a function or a set of functions as described for the modules concerned. In the same way, a hardware component corresponds to any element of a hardware assembly capable of implementing a function or a set of functions for the module concerned (integrated circuit, smart card, memory card, etc.).

[0057] The invention also relates to a server, a gateway or a terminal characterized in that it comprises a construction device as described above.

[0058] The invention also relates to a computer program comprising instructions for implementing the above method according to any of the particular embodiments described above, when said program is executed by a processor. The method can be implemented in various ways, including in hard-wired form or in software form. This program can use any programming language, and be in the form of source code, object code, or intermediate code between source code and object code, such as in a partially compiled form, or in any other desirable form.

[0059] The invention also relates to a recording medium or information medium readable by a computer, and comprising instructions of a computer program as mentioned above. The recording media mentioned above can be any entity or device capable of storing the program. For example, the medium can comprise a storage means, such as a ROM, for example a CD ROM or a microelectronic circuit ROM, or a magnetic recording means, for example a hard disk. Furthermore, the recording media can correspond to a transmissible medium such as an electrical or optical signal, which can be conveyed via an electrical or optical cable, by radio or by other means. The programs according to the invention can in particular be downloaded from a network such as the Internet.

[0060] Alternatively, the recording media may correspond to an integrated circuit in which the program is incorporated, the circuit being adapted to carry out or to be used in carrying out the method in question.

[0061] This construction device and this computer program have characteristics and advantages similar to those described previously in relation to the construction method.

[0062] 4. List of figures Other characteristics and advantages of the invention will appear more clearly on reading the following description of particular embodiments, given as simple illustrative and non-limiting examples, and the appended drawings, among which:

[0063] [Fig 1] Figure 1 illustrates an example of an environment for implementing the invention according to a particular embodiment of the invention,

[0064] [Fig 2] Figure 2 illustrates the architecture of a device suitable for implementing the construction method, according to a particular embodiment of the invention;

[0065] [Fig 3] Figure 3 illustrates the main steps of the construction method according to a particular embodiment of the invention.

[0066] 5. Description of an embodiment of the invention

[0067] Figure 1 illustrates an example of an environment for implementing the invention according to a particular embodiment. The environment represented in Figure 1 comprises at least one terminal 101 which integrates a construction device capable of implementing the construction method according to the present invention.

[0068] The process can operate permanently and autonomously upon activation of the device or following a user action.

[0069] Terminal 101 is, for example, a terminal of the smartphone type (smartphone in English), tablet, connected television, connected object, car on-board computer, personal computer, server, gateway, etc.

[0070] One or more graphics / display rendering devices (105) may be included by the terminal 101 or connected (connected wired via a VGA, HDMI, USB, etc. cable or wirelessly via WiFi®, Bluetooth®, etc. technology). This or these rendering devices may be a screen or a video projector.

[0071] According to a particular embodiment of the invention, the graphic rendering peripheral(s) can be connected to the terminal 101 via the network 102.

[0072] Similarly, one or more input / pointing devices (103a, 103b) may be included by the terminal 101 or connected (connected wired via a VGA, HDMI, USB, etc. cable or wirelessly via WiFi®, Bluetooth®, etc. technology). This or these pointing devices may be a keyboard, a mouse, a touch surface, a camera (104), a microphone or any other device capable of providing location and action data at the level of an element displayed by a display device of the terminal 101.

[0073] Figure 2 illustrates a device (S) configured to implement the construction method according to a particular embodiment of the invention. The device (S) has the conventional architecture of a computer, and notably comprises a memory MEM, a processing unit UT, equipped for example with a processor PROC, and controlled by the computer program PG stored in memory MEM. The computer program PG comprises instructions for implementing the steps of the construction method as described later in support of Figure 3, when the program is executed by the processor PROC.

[0074] At initialization, the code instructions of the computer program PG are for example loaded into a memory before being executed by the processor PROC. The processor PROC of the processing unit UT notably implements the steps of the construction method according to any one of the particular embodiments described in relation to FIG. 3 and according to the instructions of the computer program PG.

[0075] The device (S) comprises an obtaining module OBT1 capable of obtaining at least one position data item of a cursor associated with a pointing / input device. The position data item can be obtained following an action performed at the pointing device. This action can be a movement (for example a movement symbolizing a cross performed using the cursor of the pointing device, the coordinates then being able to correspond to the coordinates of the point of intersection of the two lines forming the cross), a click (the coordinates being able to correspond to the coordinates of the cursor of the pointing device at the time of the click, i.e. when pressure is exerted or released on a button of the pointing device), a long press (the coordinates being able to correspond to the coordinates of the cursor of the pointing device at the time of the long press, for example, a press lasting several seconds), etc.

[0076] The device (S) comprises an obtaining module OBT2 capable of obtaining at least one system data item from the operating system of the terminal 101. The system data item may be a command, the name of a computer process executed by the operating system of the terminal 101 or any information obtained and / or generated by the operating system of the terminal 101. The device (S) further comprises an obtaining module OBT3 capable of obtaining at least one context data item as a result of an analysis of at least one digital image of a capture of at least one part of the rendering of at least one screen of the terminal 101. The technique used for obtaining a context data item may be an optical character recognition and / or computer vision technique.

[0077] The device (S) also includes an update module MAJ capable of feeding and updating a knowledge base such as a centralized and / or distributed database or one or more files.

[0078] Figure 3 illustrates the steps of the method for constructing a knowledge base (for example a database) according to a particular embodiment of the invention. Once this knowledge base is sufficiently rich in information, an OTT application can exploit it in order to offer innovative services such as, for example, a single interface capable of controlling the computer applications installed on the electronic terminal. The OTT application can also autonomously and transversely control the functions of the applications of the electronic terminal according to predetermined computer routines (a sequence of instructions).

[0079] The method is implemented by a construction device. The construction device implementing the method is integrated into, or merged with, the user's terminal 101 (this terminal is, for example, a fixed or portable personal computer, a digital tablet, a personal digital assistant, a smartphone, a workstation, an on-board computer, etc.). In a second implementation, the construction device implementing the method is integrated into, or merged with, another electronic device that cooperates with the user's terminal. This other device is, for example, a server, a home gateway, a smartphone, a connected object, etc.

[0080] According to another embodiment, the construction device may be located in the network and / or distributed across one or more computing machines such as computers, terminals or servers.

[0081] In this particular embodiment, it is assumed that the user's terminal allows multi-windowing, i.e. the simultaneous display of several application windows on the terminal's screen. As already mentioned above, an application window is a window linked to the execution of an application by the terminal. It is also assumed that the terminal's operating system makes it possible to retrieve, and provide to the construction device that implements the present method, certain system events such as the position of a cursor of a pointing device associated with the user's terminal, the name of a computer process or a computer command that has recently been executed by the terminal.

[0082] In the first step (GET1) the method obtains cursor position data from a pointing device / peripheral (a touchpad, a mouse, a trackball, a trackpoint, a joystick or even from a camera (via an eye tracking process).

[0083] This position data may correspond to pixel coordinates or to a ratio (relative position, for example, as a percentage) of the size of the screen and / or of a graphic window when a particular system event is detected. The position data is, for example, obtained following an action performed by a user of the terminal 101 at an input device. This action may be a click performed via a mouse, the pressing of a particular key on a keyboard, a selection command spoken by the user, a particular gesture captured by a camera included or associated with the terminal 101 (for example the camera 104) or any action allowing interaction with a human-machine interface rendered by the terminal 101 via the screen 105.

[0084] During this step, the method also obtains an image representative of the content displayed by a display device (105) of the terminal 101. The image is for example obtained from software executed by the terminal 101 capable of capturing the graphic content of the display device 105 of the terminal 101.

[0085] Alternatively, the image is obtained from a third-party terminal (e.g., a camera or smartphone) positioned to capture the content displayed by a display device (105) of the terminal 101. This latter case may involve the transmission of the image, by the third-party terminal, to the terminal 101.

[0086] Alternatively, the method generates an image representative of the graphic content of the display device 105 of the terminal 101.

[0087] According to a particular embodiment of the invention, the image obtained corresponds to a part of the content rendered by the display device 105, for example, the active graphics window (the window that holds the focus). During the second step (GET2) the method obtains a system data item from the operating system of the terminal 101. This system data item may include the name of a process currently running and / or a system command interpretable by the operating system of the terminal 101 and capable of triggering an action / event at the terminal 101 (for example the launch or closing of software, the execution of a function, etc.). This system data item may also include the name of the active application (application that has the focus) as well as execution parameters such as the size of the active graphics window of the application and its position within the screen.

[0088] According to a particular embodiment of the invention, the method can obtain the system data from software and / or a third-party device.

[0089] In step (GET3) the method analyzes the image obtained during step GET1. The analysis may include optical character recognition applied to the image. Optical character recognition makes it possible to obtain a transcription of the text contained in the image in the form of a sequence of characters.

[0090] The analysis may further include recognition using a computer vision technique. This technique makes it possible to extract information from graphic elements contained within the image. For example, text recognition, object recognition, recognition of a specific element (image representing an animal, a vehicle, etc.). These techniques (optical character recognition and computer vision) may be applied to a part of the image (for example, centered according to the position data obtained during step GET1).

[0091] In a particular implementation, each text element (e.g., line) or image recognized by a recognition technique (OCR, computer vision, etc.) is ignored if a confidence score associated with the recognition is lower than a particular threshold (configuration parameter). In other words, extracted information is only taken into account if it is associated with a detection confidence score higher than a confidence value.

[0092] According to a particular embodiment of the invention, the optical character recognition and / or computer vision is / are performed by third-party software. The result(s), for example, the transcription of the text (sequence / sequence of characters and their associated position) and / or the recognized graphic elements, are sent by the third-party software to the method. According to a particular embodiment of the invention, the method performs the optical character recognition and / or computer vision at the image level.

[0093] Once the analysis is completed, the process results in a set of contextual data. This data can belong to the group including:

[0094] - a text (for example the description of a given function via a pop-up or a tooltip);

[0095] - a label (function name and / or keyboard shortcut) associated with a graphic element (button, menu, window); the size of a graphic element (button, window, menu, text, label, etc.) within the image; the position of a graphic element within the image;

[0096] - an image associated with a function of the application (for example an image of a graphic element (button, menu, etc.) capable of triggering the execution of a function following a user action (click, voice command, pressing a key on a keyboard, etc.); etc.

[0097] The label may also include the nature / type of the primary function rendered by a displayed and / or active graphical window (e.g., “email editing,” “video conferencing,” “instant messaging” window, etc.).

[0098] Of course, the preceding list of context data is not exhaustive, and other types of context data are likely to be determined during image analysis.

[0099] Note that the position can correspond to pixel coordinates or to a ratio (relative position, for example, in percentage) of the size of the graphic window that contains the graphic element. The information on the position and the size is formed by an X,Y position (for example, relative to an angle of the rectangular window) and a pair (height, width). The present invention is not limited to rectangular application windows, but applies regardless of the shape (round, oval, etc.).

[0100] During the UPDATE step, the method updates a knowledge base (KB) with the position data, the system data and the context data. When the database is sufficiently full (i.e. the learning period has been sufficiently long), it can be used by a third-party application to control the applications present on the terminal 101.

[0101] In fact, this knowledge base can be considered as a table of correspondences between each function of each application (computer application) used by the user and the associated system actions (actions executed by the electronic terminal) and the position of these functions in an application window.

[0102] Example of application: A user regularly uses videoconferencing software. The construction method is, for example, executed when the user clicks (system event within the meaning of the invention) on the icon for hanging up / leaving the videoconference. At the time of the click, the method can obtain: the position of the mouse cursor on the user's computer (position data within the meaning of the invention);

[0103] - an image of the content displayed by the user's screen corresponding to the active window of the videoconferencing software; the position of the window and / or the position of the cursor relative to the window; the name of the computer process that has just performed the action. The name of the process may correspond to the name of the software / active window. It is obtained, for example, by monitoring the use of the processor and / or memory (variation / peak consumption). Note that the name of the active window may also be obtained in response to the execution of a command sent to the terminal's operating system.

[0104] The method then performs an analysis of the image using an optical character recognition and / or computer vision technique. The method can obtain the following context data as a result:

[0105] - a label for the active graphic window (which can correspond to the name of the application, a version number, a document title, etc.); the size of the active application window (length and width) in pixels;

[0106] - a label located near the cursor position. In this particular case, the label corresponds to "quit". Note that the label corresponds to the function invoked by the user when they clicked (i.e., quit the videoconference). The label can also correspond to a keyboard shortcut. In our case, the shortcut can correspond to "Ctrl + F4"; - an image / icon located near the cursor position (i.e., associated with the function invoked by the user). In our case, it can be a clickable pictogram symbolizing a cross allowing you to quit the videoconference);

[0107] - a label indicating the type / nature of the active window. In our case, the label could correspond to “videoconference”.

[0108] Note that the obtained data can be ignored if a confidence score associated with optical character recognition and / or computer vision is lower than a particular threshold (configuration parameter). In other words, extracted information is only taken into account if it is associated with a detection confidence score higher than a confidence value.

[0109] The process then updates or adds all or part of this information to the BDD knowledge base (distributed or non-distributed database). A database record can have a format of the type:

[0110] {"function", "application name", "X position", "Y position"}

[0111] Or {"function", "application name", "X position", "Y position", "keyboard shortcut"};

[0112] Or {“function”, “application name”, “X position”, “Y position”, “type”}•

[0113] The X and Y positions can correspond to the relative coordinates (for example in percentage) of the cursor within the active application window.

[0114] In the case described above the recording may correspond to:

[0115] {"leave", "Microsoft Teams", "left:0.85", "top:0.09", "videoconference"}

[0116] Of course, before adding this record, the process checks whether a record of the same type does not exist in the database. This check consists of searching the database for a record with an identical primary key. In our case, the primary key can correspond to: {“function”, “application name”}; to {“image / icon of the function”, “application name”}; or to {“function”, “application name”, “type”}. If such a record exists, the process compares the data obtained with those present in the record. If the data match, no processing is carried out by the process. Otherwise, the process updates the data of the found record with the data previously obtained from the terminal’s operating system and the analysis of the image.

[0117] When the database is sufficiently full, it can be used by a third-party application to control the applications present on the user's computer.

[0118] Thus, when the third-party application requests to leave a videoconference (for example following a user query), it retrieves the relevant data (for example the position of the clickable graphic element corresponding to the “leave” function, i.e. the X and Y coordinates) from the knowledge base via, for example, a search having the name of the function and the name of the software as the primary key. The third-party application (OTT) then replays the click (simulates the user’s click) at the position (X,Y) obtained previously.

[0119] Obviously this assumes that the function invoked by the user through the third-party application (OTT) is displayed on the screen of the terminal 101.

[0120] According to a particular embodiment of the invention, the method can also replay the keyboard shortcut corresponding to the function invoked by the user.

[0121] According to a particular embodiment of the invention, the method can, during the step of updating the knowledge base, add an identifier associated with the active application window. This identifier is for example obtained via the application of a cryptographic function at the level of the image obtained. The cryptographic function can be a hash function allowing a digital fingerprint of the image to be obtained. This embodiment makes it possible, when the primary key includes this identifier, to distinguish the same function (for example "exit") which would be present within two different graphic windows of the same application.

[0122] This identifier may also include data obtained from the operating system of the terminal 101, such as the name and / or the access path (tree structure of directory(ies) / computer file(s)) to the executable of the active application window. Note that when the identifier of the window corresponds to a digital fingerprint, this assumes that the content of the window does not vary.

[0123] According to a particular embodiment of the invention, the method can obtain from the operating system the system command generated following the user's mouse click. This command can then be included in the recording obtained by the third-party application (OTT). In the case described above, the recording can correspond to:

[0124] {"function", "application name", "X position", "Y position", "command"}

[0125] The command can then be executed by the computer's operating system at the request of the third-party application (i.e., the user).

[0126] Thus, if the "quit" functionality is available via a button in the human-machine interface but also via an "option / quit" submenu, the system command asking the application to quit the videoconference will be the same (regardless of whether the user clicked on the button or went through the menu / submenu). Furthermore, and unlike the solution described previously based on a replay of a mouse click or a keyboard shortcut, it is not necessary for the function invoked by the third-party application (OTT) to be displayed on the screen of the terminal 101.

Claims

CLAIMS 1. Method for constructing a knowledge base, said method being implemented by a construction device during use of an electronic terminal (101), and characterized in that the method comprises: when a system event on said electronic terminal is detected - a first step of obtaining (GET1) at least one item of cursor position data, said cursor being associated with at least one pointing device (103a, 103b) of said electronic terminal, and at least one digital image of a capture of at least one part of the rendering of at least one screen of said electronic terminal; - a second step of obtaining (GET2) at least one system data item from said electronic terminal; - a third step of obtaining (GET3) at least one piece of context data based on an analysis of all or part of said digital image; - an update step (UPDATE) of said knowledge base with said at least one position, system and context data.

2. Method according to claim 1 characterized in that said system event is generated by at least one pointing device (103 a, 103 b) associated with said electronic terminal (101).

3. Method according to claim 1 characterized in that the capture relates to an active application window.

4. Method according to claim 1 characterized in that obtaining said at least one context data item is carried out via an optical character recognition technique and / or a computer vision technique.

5. Method according to claim 1 characterized in that the updating step is conditioned by the value of a confidence score associated with said at least one piece of context data.

6. Method according to claim 1 characterized in that said at least one system data item comprises at least one computer command executable by the operating system of said electronic terminal.

7. Method according to claim 1 characterized in that said context data belongs to the group comprising at least: - a label; - a graphics window size; - a position within said screen; - a text; - an image.

8. Device for constructing a knowledge base implemented during use of an electronic terminal, and characterized in that the device comprises: - a first module (OBT1) for obtaining at least one item of cursor position data, said cursor being associated with said at least one pointing device (103a, 103b) of said electronic terminal, and at least one digital image of a capture of at least one part of the rendering of at least one screen of said electronic terminal; - a second module (OBT2) for obtaining at least one system data item from said electronic terminal; - a third module (OBT3) for obtaining at least one piece of context data based on an analysis of all or part of said digital image; - an update module (MAJ) of said knowledge base with said at least one position, system and context data.

9. Server, gateway or terminal characterized in that it comprises a construction device according to claim 8.

10. Computer program comprising the instructions for executing the construction method according to any one of claims 1 to 7, when the program is executed by a processor.