Information processing method and device, electronic equipment, storage medium and program product

By displaying additional information of image information in response to user triggering operations in the human-computer interaction interface, the problem of difficult to directly present object information in video and image is solved, and fast and accurate information acquisition and efficient human-computer interaction are achieved.

CN120255747APending Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410004524.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-02
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, the information and relationships of multiple objects in video and image information cannot be directly presented, resulting in limited user understanding and additional search and acquisition, resulting in low information acquisition speed and low accuracy.

Method used

The image information is displayed in the human-computer interaction interface, and the additional information of the object is displayed in response to the user's trigger operation. The degree of detail of the additional information is related to the operation parameters, including the object association relationship and attribute information.

Benefits of technology

The information representation ability of image information is improved, and users can quickly and accurately obtain additional information, improving human-computer interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120255747A_ABST
    Figure CN120255747A_ABST
Patent Text Reader

Abstract

The invention provides an information processing method and device, electronic equipment, a computer readable storage medium and a computer program product. The method can be applied to a vehicle-mounted scene, and the method comprises the steps that image information is displayed in a human-computer interaction interface, and the image information comprises at least one object; in response to a first trigger operation for the human-computer interaction interface, displaying additional information for describing the object on the image information; wherein the detailed degree of the additional information is related to the operation parameters of the first trigger operation. According to the invention, the information representation capability of the image information and the man-machine interaction efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to video processing technologies, and in particular, to an information processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] In related technologies, rich content can be presented to users through video and image information such as videos and images. Multiple objects are often included in videos and images, and the information of each object and the relationships between objects usually cannot be directly presented in the currently played video or the currently displayed image, resulting in limited understanding of image information by users or even confusion.

[0003] In related technologies, users can retrieve object relationship diagrams and detailed object introductions through a search engine. However, this requires additional search processing by users, resulting in a low information acquisition speed, and the relevance between the obtained object relationship diagrams and detailed object introductions and the image information currently viewed by users may be weak, resulting in low information acquisition accuracy. Summary of the Invention

[0004] Embodiments of the present application provide an information processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can improve the information representation ability of image information while obtaining additional information faster and more accurately, thereby improving the human-computer interaction efficiency.

[0005] The technical solution of the embodiments of the present application is implemented as follows:

[0006] Embodiments of the present application provide an information processing method, including:

[0007] Displaying image information in a human-computer interaction interface, where the image information includes at least one object;

[0008] Responding to a first trigger operation on the human-computer interaction interface, and displaying additional information for describing the object on the image information;

[0009] Wherein, the detail level of the additional information is related to the operation parameter of the first trigger operation.

[0010] Embodiments of the present application provide an information processing apparatus, including:

[0011] A display module, configured to display image information in a human-computer interaction interface, where the image information includes at least one object;

[0012] An additional information module, configured to display additional information for describing the object on the image information in response to a first trigger operation on the human-computer interaction interface; wherein, the detail level of the additional information is related to the operation parameters of the first trigger operation.

[0013] In some embodiments, the additional information module is further configured to: display additional information for describing the object on the image information; wherein, the additional information includes at least one of the following: the association relationship between multiple objects, the attribute information of the object, and the attribute information of the object includes at least one of the following: the role name of the object, the real name of the object.

[0014] In some embodiments, the additional information module is further configured to: display the object in a marked state in the image information, and display the attribute information of the object at the associated position of the object in the marked state; for two objects with an association relationship, display a connection line for connecting the two objects, and display the association relationship between the two objects at the associated position of the connection line.

[0015] In some embodiments, the additional information module is further configured to: display an additional information display mode control, wherein the additional information display mode control is used to switch between an on state and an off state when triggered, the on state indicates that the additional information display mode is turned on, and the off state indicates that the additional information display mode is turned off; in response to the additional information display mode control being in the on state and a first trigger operation on the human-computer interaction interface, display additional information for describing the object on the image information.

[0016] In some embodiments, the additional information module is further configured to: start timing from the display of the additional information after displaying the additional information for describing the object on the image information; when the timing reaches a first time threshold, display the additional information display mode control in the off state, and at the same time stop displaying all additional information, or stop displaying all additional information in sequence.

[0017] In some embodiments, the image information is a video frame in a video, and the additional information module is further configured to: display the additional information on the first video frame currently being played in the video; play to the second video frame of the video, and the second video frame is a video frame in the video that is after the first video frame; when the scene similarity between the second video frame currently being played in the video and the first video frame is less than the scene similarity threshold, stop displaying all additional information at the same time.

[0018] In some embodiments, the additional information module is further configured to: stop displaying all the additional information in the following order: the order from the lowest to the highest importance level of the additional information; the order from the lowest to the highest importance level of the objects to which the additional information belongs.

[0019] In some embodiments, the additional information module is further configured to perform any one of the following processes: in response to a first trigger operation on a target object in the image information, display the additional information of the target object on the image information, where the target object is from the at least one object; in response to a first trigger operation on a target text in the image information, display the additional information of the target object on the image information, where the target object is the object specified by the target text and the target object is from the at least one object; in response to a first trigger operation on any position in the image information, display the additional information of the target object on the image information, where the target object is the object with the highest importance level among all the objects in the image information.

[0020] In some embodiments, the additional information module is further configured to, before responding to a first trigger operation on the human-computer interaction interface, perform any one of the following processes: in response to a pressing operation on the image information, identify the pressing operation as the first trigger operation on the human-computer interaction interface; in response to a clicking operation on the image information, identify the clicking operation as the first trigger operation on the human-computer interaction interface; display an additional information display control, and in response to a control trigger operation on the additional information display control, identify the control trigger operation as the first trigger operation on the human-computer interaction interface.

[0021] In some embodiments, when the first trigger operation is a non-control trigger operation on the human-computer interaction interface, the operation parameters include at least one of the following: the contact time of the non-control trigger operation, the contact force of the non-control trigger operation, the contact frequency of the non-control trigger operation; when the first trigger operation is a control trigger operation on the additional information display control, the operation parameters include: the control parameters of the additional information display control caused by the control trigger operation.

[0022] In some embodiments, the additional information module is further configured to: in response to meeting the condition for automatically displaying additional information, display additional information for describing the object on the image information; the condition for automatically displaying additional information includes at least one of the following: the number of objects included in the image information exceeds an object number threshold; there is barrage information displayed on the image information, and the barrage information indicates that the audience of the image information has a need to obtain additional information about the at least one object; the subtitle of the image information includes content related to the at least one object.

[0023] An embodiment of the present application provides an electronic device, including:

[0024] A memory for storing computer-executable instructions;

[0025] A processor, when executing the computer-executable instructions stored in the memory, implements the information processing method provided by the embodiment of the present application.

[0026] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing a processor to implement the information processing method provided by the embodiment of the present application when executed.

[0027] An embodiment of the application provides a computer program product including computer-executable instructions that, when executed by a processor, implement the information processing method provided by the embodiment of the present application.

[0028] The embodiment of the present application has the following beneficial effects:

[0029] An image information including at least one object is displayed in a human-computer interaction interface, where a video including at least one object can be played or an image including at least one object can be displayed. In response to a first trigger operation on the image information, additional information of the object is displayed on the image information. This is equivalent to directly displaying the additional information of the object through the user's trigger operation, which helps the user efficiently and comprehensively perceive the image information, improves the information representation ability of the image information, and enables the user to obtain the additional information faster and more accurately, thereby improving the human-computer interaction efficiency. Moreover, the detail level of the additional information is related to the operation parameters of the first trigger operation, which is equivalent to directly controlling the detail level of the additional information through the operation parameters of the first trigger operation, meeting the user's demand for the detail level of the additional information and further improving the human-computer interaction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figures 1A - 1C is a schematic diagram of an interface of an information processing method provided in the related art;

[0031] Figure 2 is a schematic diagram of the structure of an information processing system provided by the embodiment of the present application;

[0032] Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0033] Figures 4A - 4C is a schematic flowchart of an information processing method provided by an embodiment of the present application;

[0034] Figures 5A - 5C is a schematic interface diagram of an information processing method provided by an embodiment of the present application;

[0035] Figure 6 is a schematic implementation logic diagram of an information processing method provided by an embodiment of the present application. Detailed implementation manners

[0036] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0037] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0038] In the following description, the terms "first / second / third" are merely used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0040] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.

[0041] 1) Force Touch, which triggers different functions by applying different degrees of pressure on the touch screen. It can identify different gestures according to the pressure applied by the user on the touch screen, such as light press, hard press, or long press. In the embodiments of the present application, the pressure sensor of the touch screen is mainly used to detect the pressure level applied by the user's finger.

[0042] 2) Video understanding, which automatically identifies and analyzes video content through intelligent analysis technology. Image classification is the basis of video understanding. In the embodiments of the present application, a recurrent neural network is used for video understanding, and the video is regarded as time-series data composed of each frame of image for understanding.

[0043] In the related art, rich content can be presented to users through videos and other image information. Refer to Figures 1A - 1B , there are often multiple objects in videos and images. When users watch video content, their understanding of image information is limited or even confusing, and they need to manually drag back to the initial position to understand many plot information including character names and relationships, resulting in a poor viewing experience.

[0044] Refer to Figure 1C , in the related art, users can retrieve object relationship diagrams and object details through a search engine, but this requires users to perform additional search processing. Users cannot watch videos and detailed introductions at the same time, resulting in a low information acquisition speed, and the relevance between the obtained object relationship diagrams and object details and the image information currently viewed by users may be weak, resulting in a low information acquisition accuracy.

[0045] The embodiments of the present application provide an information processing method, device, electronic device, computer-readable storage medium, and computer program product, which can improve the diversity and efficiency of live interaction.

[0046] The following describes an exemplary application of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as a terminal or a server.

[0047] Refer to Figure 2 , Figure 2 is a schematic diagram of the application mode of the information processing method provided by the embodiments of the present application; for example, Figure 2 involves a server 200, a network 300, and a terminal 400. The terminal is connected to the server 200 through the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.

[0048] In some embodiments, the server 200 can be a server corresponding to an application program. For example, if the application program is a video software installed in the terminal, the server 200 is a server of the video platform.

[0049] In response to the terminal 400 receiving a user's play operation on a target video, the terminal 400 sends a play request to the server 200. The server 200 transmits video data to the terminal 400, and the terminal 400 plays the video data. In response to the terminal 400 receiving a first trigger operation on the video data, the terminal 400 sends a display request to the server 200. Here, the display request carries the level of detail corresponding to the intensity of the first trigger operation. The server 200 calls the additional information of the object based on the level of detail and returns the additional information of the object that meets the level of detail to the terminal 400 to display the additional information of the object on the video.

[0050] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart TV, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, and there is no limitation in the embodiments of the present application.

[0051] In some embodiments, the terminal may implement the information processing method provided in the embodiments of the present application by running a computer program. For example, the computer program may be a native program or software module in an operating system; it may be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a video APP; it may also be a small program, that is, a program that only needs to be downloaded to the browser environment to run; it may also be a small program that can be embedded in any APP. In short, the above computer program may be any form of application program, module or plug-in.

[0052] See Figure 3 , Figure 3 is a schematic structural diagram of an electronic device provided in the embodiments of the present application. The electronic device is a terminal. Figure 3 The terminal shown includes: at least one processor, a memory 450, at least one network interface 420, and a user interface 430. Each component in the terminal is coupled together through a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between these components. The bus system 440 includes not only a data bus, but also a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 3 all kinds of buses are labeled as the bus system 440.

[0053] The processor can be an integrated circuit chip with the ability to process signals, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0054] The user interface 430 includes one or more output devices 431 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 430 also includes one or more input devices 432, including user interface components that facilitate user input, such as a keyboard, a mouse, a microphone, a touch screen display, a camera, other input buttons, and controls.

[0055] The memory 450 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 450 optionally includes one or more storage devices that are physically remote from the processor.

[0056] The memory 450 includes volatile memory or non-volatile memory, and may also include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.

[0057] In some embodiments, the memory 450 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are described below by way of example.

[0058] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0059] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wi-Fi (Wireless Fidelity), and USB (Universal Serial Bus), etc.;

[0060] A presentation module 453 for enabling presentation of information (e.g., a user interface for operating a peripheral device and displaying content and information) via one or more output devices 431 associated with the user interface 430 (e.g., a display screen, a speaker, etc.).

[0061] An input processing module 454 for detecting one or more user inputs or interactions from one of one or more input devices 432 and translating the detected inputs or interactions.

[0062] In some embodiments, the device provided by the embodiments of the present application may be implemented in software. Figure 3 An information processing device 455 stored in the memory 450 is shown, which may be software in the form of programs and plugins, etc., including the following software modules: a display module 4551 and an additional information module 4552. These modules are logical, and thus can be arbitrarily combined or further split according to the implemented functions. The functions of each module will be described below.

[0063] The information processing method provided by the embodiments of the present application will be described in combination with the exemplary applications and implementations of the terminal provided by the embodiments of the present application.

[0064] Next, the information processing method provided by the embodiments of the present application will be described. As before, the electronic device implementing the information processing method of the embodiments of the present application may be a terminal or a server. Taking the terminal as an example for description. Therefore, the execution subject of each step will not be repeated hereinafter.

[0065] See Figure 4A , Figure 4A is a schematic flowchart of the information processing method provided by the embodiments of the present application, which will be described in combination with Figure 4A the steps shown.

[0066] In step 101, image information is displayed in the human-computer interaction interface.

[0067] As an example, the image information includes at least one object. Here, the image information may be a single image or a video frame in a video. If the image information is a video frame in a video, the displayed image information may be a played video frame or a paused video frame. Here, the object is a person in the image or video, which may be an animated character or a character in a TV drama. See Figure 5A , and an object 504A is shown in the image information.

[0068] In step 102, in response to a first trigger operation on the human-computer interaction interface, additional information for describing the object is displayed on the image information.

[0069] As an example, the level of detail of the additional information is related to the operation parameters of the first trigger operation. Examples of the level of detail of the additional information are as follows. The additional information may include a name and an age, or the additional information may only include a name. When the additional information includes a name and an age, the level of detail is higher than when the additional information only includes a name.

[0070] As an example, when the first trigger operation is a non-control trigger operation on the human-computer interaction interface (for example, the non-control trigger operation here may be a click operation or a press operation), the operation parameters include at least one of the following: the contact time of the non-control trigger operation (the level of detail may be positively correlated with the contact time), the contact force of the non-control trigger operation (the level of detail may be positively correlated with the contact force), the contact frequency of the non-control trigger operation (the contact frequency may be positively correlated with the contact force); when the first trigger operation is a control trigger operation on the additional information display control, the operation parameters include: the control parameters of the additional information display control caused by the control trigger operation. Here, the form of the additional information display control may be similar to a volume control. Through the control trigger operation, the progress bar in the additional information display control is presented at different heights, and here the height is the control parameter.

[0071] Through the embodiments of the present application, the level of detail of the additional information can be controlled while triggering the display of the additional information. Therefore, it is equivalent to realizing the reuse of the first trigger operation and improving the human-computer interaction efficiency.

[0072] In some embodiments, the additional information includes at least one of the following: the association relationship between multiple objects, the attribute information of the object. The attribute information of the object includes at least one of the following: the role name of the object, the real name of the object.

[0073] As an example, the additional information may be the real name of the actor of the object, the role name of the object, the identity of the object in the play, the relationship between the roles played by the object. Here, the relationship may be a mother-daughter relationship, a husband-wife relationship, etc.

[0074] Through the embodiments of the present application, the additional information can be displayed in a diversified and comprehensive manner, and the relationship between each additional object can be shown, so that the information is integrated and displayed.

[0075] In some embodiments, in step 102, the additional information for describing the object is displayed on the image information, the object in the marked state is displayed in the image information, and for the object in the marked state, the attribute information of the object is displayed at the associated position of the object; for two objects with an association relationship, a connection line for connecting the two objects is displayed, and the association relationship between the two objects is displayed at the associated position of the connection line.

[0076] As an example, refer to Figure 5B , in the image information, there are two objects (object A and object B) in the marked state (the heads are circled). The attribute information of object A can be displayed at the associated position of object A, and the attribute information of object B can be displayed at the associated position of object B. Here, the associated position can be the position directly above the head of the object. The attribute information displayed here is the name (Zhang San) and age (36 years old) of the role played by the object. Figure 5B also shows a connection line between object A and object B, and the association relationship between object A and object B (starting a business in partnership) is displayed at the associated position of the connection line. Here, the associated position can be the position directly above the connection line.

[0077] Through the embodiments of the present application, the additional information can be closely associated and displayed with the object, that is, in a visual way, the user can clearly know which user the additional information is used to describe, improving the human-computer interaction efficiency and information representation ability.

[0078] In some embodiments, refer to Figure 4B , in step 102, in response to a first trigger operation on the human-computer interaction interface, the additional information for describing the object is displayed on the image information, which can be implemented through Figure 4B the steps 1021 to 1022 shown.

[0079] In step 1021, an additional information display mode control is displayed, where the additional information display mode control is used to switch between the on state and the off state when triggered. The on state represents the opening of the additional information display mode, and the off state represents the closing of the additional information display mode;

[0080] In step 1022, in response to the additional information display mode control being in the on state and a first trigger operation on the human-computer interaction interface, the additional information for describing the object is displayed on the image information.

[0081] As an example, refer to Figure 5A , in the human-computer interaction interface 501A, an additional information display mode control 502A is displayed, and Figure 5AThe additional information display mode control 502A shown is in the on state. Through a triggering operation on the additional information display mode control 502A, the additional information display mode control 502A can be switched from the on state to the off state, or from the off state to the on state. When the additional information display mode control is in the on state, in response to a first triggering operation by the user (the user viewing the image information) on the human-machine interface, the additional information 503A shown is displayed on the image information. When the additional information display mode control is in the off state, in response to a first triggering operation by the user (the user viewing the image information) on the human-machine interface, no additional information is displayed on the image information. Figure 5A By the embodiments of the present application, conflicts between the first triggering operation and current existing gestures can be avoided, thereby increasing the possibility of gesture reuse, and mis-triggering can be avoided, improving the human-machine interaction efficiency.

[0082] In some embodiments, after the additional information for describing the object is displayed on the image information, timing starts from the display of the additional information; when the timing reaches a first time threshold, the additional information display mode control is displayed in the off state, and at the same time, the display of all additional information is stopped, or the display of all additional information is stopped sequentially.

[0083] As an example, referring to

[0084] When the additional information is displayed on the image information, timing starts. If no first triggering operation is received within the first time threshold, then the additional information display mode control is displayed in the off state, that is, the additional information display mode control 501C is in the off state, and no additional information is displayed. Figure 5C

[0085] In some embodiments, after the additional information for describing the object is displayed on the image information, timing starts from the display of the additional information; when the timing reaches a second time threshold, the display of all additional information is stopped at the same time, or the display of all additional information is stopped sequentially.

[0086] Figure 5C As an example, referring to When the additional information is displayed on the image information, timing starts. If no first triggering operation is received within the second time threshold, the additional information is no longer displayed.

[0087] By the embodiments of the present application, excessive occupation of screen resources by additional information can be avoided. The additional information is only displayed for a fixed duration, and the additional information display mode is only turned on for a fixed duration, so that the normal interaction of the user with the human-machine interface can be restored.

[0088]

[0088] In some embodiments, the image information is a video frame in a video. Displaying additional information for describing the object on the image information in step 102 can be achieved through the following technical solutions: Display the additional information on the first video frame currently being played in the video; stopping the display of all additional information simultaneously can be achieved through the following technical solutions: Play to the second video frame of the video, where the second video frame is a video frame in the video that comes after the first video frame; when the scene similarity between the second video frame currently being played in the video and the first video frame is less than the scene similarity threshold, stop the display of all additional information simultaneously.

[0089] As an example, when the image information is a video frame in a video, in fact, the additional information is displayed on the first video frame being played in real time. Correspondingly, as the video continues to play, the second video frame that comes after the first video frame will be played. If the scene similarity between the second video frame played later and the first video frame is less than the scene similarity threshold, for example, the first video frame presents a conversation scene between object A and object B in an office, and the video frames after the first video frame may still continue to present the conversation scene between object A and object B in the office, then the additional information of object A and object B can continue to be displayed. However, if a scene transformation event occurs, such as switching to a scene where object C and object D are walking in the park, then the additional information of object A and object B is no longer useful, so it is necessary to stop displaying all the additional information of object A and object B. The scene similarity here can be calculated through an image understanding algorithm, calculating the vector similarity between the image feature vector of the first video frame and the image feature vector of the subsequent second video frame, and using the vector similarity to represent the scene similarity.

[0090] Through the embodiments of the present application, it is possible to automatically stop displaying all additional information when the scene changes, avoiding the situation where the additional information does not match the image information, so that the user needs to manually close the additional information later, and the human-computer interaction efficiency can be improved.

[0091] In some embodiments, the above-mentioned sequential stopping of the display of all additional information can be achieved through the following technical solutions: Stop displaying all additional information in the following order: The order from low to high of the importance level of the additional information; the order from low to high of the importance level of the object to which the additional information belongs.

[0092] As an example, in addition to stopping the display of all additional information simultaneously as described above, the display of all additional information can also be stopped sequentially. For example, first stop the display of some additional information and then stop the display of all additional information. The display of additional information can be stopped sequentially in ascending order of the importance of the additional information. For example, the importance of a character's name is higher than that of the character's age. Therefore, the display of the character's age is stopped first, and then the display of the character's name is stopped. The importance of the objects to which the additional information belongs is in ascending order. For example, if object A belongs to the protagonist and object B belongs to the supporting role, then the display of the additional information of object B is stopped first, and then the display of the additional information of object A is stopped.

[0093] Through the embodiments of the present application, the display of additional information can be stopped sequentially in order to provide a transitional visual effect and avoid the uncomfortable experience caused by the sudden disappearance of additional information.

[0094] In some embodiments, referring to Figure 4C , in step 102, in response to a first trigger operation on the human-computer interaction interface, the display of additional information for describing the object on the image information can be implemented through any one of steps 1023 to 1025 shown in Figure 4C .

[0095] In step 1023, in response to a first trigger operation on the target object in the image information, the additional information of the target object is displayed on the image information, where the target object is from the at least one object.

[0096] As an example, referring to Figure 5A , target object A is displayed in the human-computer interaction interface 501A. In response to a first trigger operation on target object A, the additional information (Zhang San) of target object A is displayed on the image information.

[0097] In step 1024, in response to a first trigger operation on the target text in the image information, the additional information of the target object is displayed on the image information, where the target object is the object specified by the target text and the target object is from the at least one object.

[0098] As an example, referring to Figure 5A , the subtitle "She started a business in partnership with Zhang San" is displayed in the human-computer interaction interface 501A. The target text is the text "Zhang San" for specifying the target object. Therefore, in response to a first trigger operation on "Zhang San", the additional information "Zhang San" of target object A is displayed on the image information.

[0099] In step 1025, in response to a first trigger operation for any position in the image information, additional information of a target object is displayed on the image information, where the target object is the object with the highest importance among all objects in the image information.

[0100] As an example, referring to Figure 5A , the human-computer interaction interface 501A includes multiple objects. If the first trigger operation has no directivity, the additional information of the object with the highest importance among the multiple objects can be displayed here. The importance of the protagonist is higher than that of the supporting role. Here, the importance can be positively correlated with the amount of lines of the role played by the object or positively correlated with the number of appearances of the role played by the object.

[0101] Through the embodiments of the present application, additional information can be presented in a directed manner, enabling the user to independently control which additional information is presented and improving the human-computer interaction efficiency.

[0102] In some embodiments, before responding to the first trigger operation for the human-computer interaction interface, any one of the following processes is performed: in response to a pressing operation for the image information, the pressing operation is recognized as the first trigger operation for the human-computer interaction interface; in response to a clicking operation for the image information, the clicking operation is recognized as the first trigger operation for the human-computer interaction interface; an additional information display control is displayed, and in response to a control trigger operation for the additional information display control, the control trigger operation is recognized as the first trigger operation for the human-computer interaction interface.

[0103] As an example, the first trigger operation here can be a pressing operation, a clicking operation, or a control trigger operation for the additional information display control. These can be pre-configured, and these human-computer interaction operations can be recognized as the first trigger operation to trigger the display of additional information.

[0104] In some embodiments, in response to meeting the additional information automatic display condition, additional information for describing the object is displayed on the image information; the additional information automatic display condition includes at least one of the following: the number of objects included in the image information exceeds an object number threshold; there is barrage information displayed on the image information, and the barrage information indicates that the audience of the image information has a need to obtain the additional information of the at least one object; the subtitle of the image information includes content related to the at least one object.

[0105] As an example, in addition to the manual trigger method, the display of additional information can also be triggered by an automatic trigger method. For example, when the conditions for automatically displaying additional information are met, the additional information can be automatically displayed. The conditions for automatically displaying additional information include at least one of the following: the number of objects included in the image information exceeds an object number threshold (when there are a large number of objects in the image information, viewers may not be able to distinguish the identity of each object, so additional information needs to be automatically displayed); there is barrage information displayed on the image information, and the barrage information indicates that the viewers of the image information have a need to obtain additional information about the at least one object (when there is barrage information displayed on the image information, the barrage information can be posted by users watching the video information. For example, the barrage information is "I can't even tell who is who anymore". At this time, the barrage information indicates that the viewers of the image information have a need to obtain additional information); the subtitles of the image information include content related to the at least one object (for example, referring to Figure 5A , the subtitle information "She started a business in partnership with Zhang San" includes content related to the object Zhang San. Therefore, additional information can be automatically displayed to help viewers understand the image information).

[0106] By means of the embodiments of the present application, additional information can be automatically displayed, improving the efficiency of human-computer interaction.

[0107] An image information including at least one object is displayed in a human-computer interaction interface. Here, a video including at least one object can be played or an image including at least one object can be displayed. In response to a first trigger operation on the image information, additional information about the object is displayed on the image information. Equivalent to directly displaying the additional information about the object through the user's trigger operation, it helps the user to efficiently and comprehensively perceive the image information, improves the information representation ability of the image information, and at the same time, obtains the additional information faster and more accurately, thereby improving the efficiency of human-computer interaction. And the detail level of the additional information is related to the operation parameters of the first trigger operation. Equivalent to directly controlling the detail level of the additional information through the operation parameters of the first trigger operation, it can meet the user's demand for the detail level of the additional information and further improve the efficiency of human-computer interaction.

[0108] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0109] In a video playback scenario, in response to the terminal 400 receiving a user's playback operation for a target video, the terminal 400 sends a playback request to the server 200. The server 200 transmits video data to the terminal 400, and the terminal 400 plays the video data. In response to the terminal 400 receiving a first trigger operation for the video data, the terminal 400 sends a display request to the server 200. The display request carries the level of detail corresponding to the intensity of the first trigger operation. The server 200 calls the additional information of the object based on the level of detail and returns the additional information of the object that meets the level of detail to the terminal 400 to display the additional information of the object on the video.

[0110] The embodiments of the present application provide a method for triggering the display of information related to video characters based on the user's finger pressure. Specifically, in response to the user's video playback operation, the video is played on the terminal, and the server analyzes the additional information that the user may need during the video viewing process (such as the real name of the actor, the name of the character, the identity of the character in the play, the relationship between the characters, etc.). In response to the user's trigger operation for the video, according to different pressure levels of the trigger operation, different degrees of information display can be mapped. When the pressure of the trigger operation is relatively large, more comprehensive additional information is displayed, and the information follows and is bound to the avatar of the character in the video in the form of a floating layer. When the pressure of the trigger operation is relatively small, less additional information is displayed.

[0111] See Figure 5A , a video is being played in the human-computer interaction interface 501A, and an additional information display mode control 502A is displayed in the human-computer interaction interface 501A. In response to the user's trigger operation for the additional information display mode control 502A, when the additional information display mode control 502A is in the on state, in response to a light press operation on the human-computer interaction interface 501A, a small amount of additional information 503A appears in the human-computer interaction interface. See Figure 5B , in response to a heavy press operation on the human-computer interaction interface 501B, a large amount of additional information 502B appears in the human-computer interaction interface. See Figure 5C , to prevent conflict with the pause gesture, this function will automatically turn off 30 seconds after the user uses this function, that is, the additional information display mode control 501C is in the off state, and the display of additional information stops.

[0112] In some embodiments, see Figure 6 , Figure 6The flowchart shows the information processing method provided by the embodiments of the present application. In step 601, an operation for the user to open a video is received. In step 602, the operation data is sent to the server. In step 603, the server analyzes and processes the video. In step 604, the server generates additional information based on the video content and intelligent role-related information and waits for a call instruction. In step 605, the server sends a rendering polling instruction to the terminal. In step 606, an operation for the user to enable the additional information mode is received. In step 607, an operation for the user to press the screen is received. In step 608, a rendering instruction is sent to the server. In step 609, the server obtains the additional information to be displayed according to the pressing force. In step 610, the server sends the additional information back to the terminal. In step 611, the additional information is presented on the video.

[0113] The server creates additional information based on the video selected by the user. The specific process is as follows:

[0114] 1. Identify the characters in the video, perform face detection and recognition using a face recognition library, and determine the identity of the characters by comparing with a known face feature library, such as face features extracted using the FaceNet or VGGFace model.

[0115] 2. Create additional information. The additional information will include the names of the characters and the relationships between the characters. These information can be obtained through bullet screens or subtitles, and the obtained information is stored in the form of file tags. Here, the stored information needs to be divided into two categories, that is, multiple information levels are designed, and each level contains different levels of detailed information. For example, a lower pressure level can display the name of the person, while a higher pressure level can display more information, such as the profile of the person and related videos. Add text tags around the people in the video to display the name, identity or other relevant information of the characters. This can be achieved by using the drawing function of the graphics library on the video frame, such as cv2.putText().

[0116] When presenting the information, realize character recognition and real-time information tracking display in the video. The specific process is as follows:

[0117] 1. Video object detection and tracking. Use a deep learning model for object detection to identify the people in the video frame. Use relevant object tracking algorithms, such as the multi-object tracking algorithm based on the Kalman filter, to track the detected people.

[0118] 2. Face recognition and feature extraction. Use a face recognition library for face detection and recognition, and determine the identity of the characters by comparing with a known face feature library, such as face features extracted using the FaceNet or VGGFace model.

[0119] 3. Real-time information display. In the interface of the video player or application, use the graphics library to display video frames in real time. According to the position and movement of the characters, display the additional information in real time at the corresponding positions in the video. You can use the drawing functions provided by the graphics library to draw text on the video frames.

[0120] A code example of the above process is as follows:

[0121]

[0122]

[0123] Display the pre-rendered information content according to the pressure on the user's touch screen. Determine whether the current user has enabled the "Additional Information" mode. If not, the user's touch on the screen is the original interaction method (i.e., pause), and a long press is the original interaction method (i.e., speed up playback).

[0124] If the user has enabled the "Additional Information" mode, you can use the API or SDK provided by the hardware manufacturer to obtain the values of the pressure sensor. According to the value range of the pressure sensor, map it to different pressure levels. The specific process can be seen in the following code:

[0125]

[0126]

[0127] Then call the pre-rendered information through the following custom mapping function. See the following code:

[0128]

[0129] If the pressure_level is zero for up to 30 seconds, it means that the user does not need additional information, so the system automatically closes the "Additional Information" mode.

[0130] The embodiment of the present application provides a method for triggering the display of information related to video characters based on the user's finger pressure. Specifically, in response to the user's video playback operation, play the video on the terminal, and analyze the additional information (such as the real name of the actor, the name of the character, the identity of the character in the play, the relationship between characters, etc.) that the user may need during the video viewing process through the server. In response to the user's trigger operation on the video, according to the different pressure levels of the trigger operation, it can be mapped to different degrees of information display. When the pressure of the trigger operation is large, more comprehensive additional information is displayed, and the information follows and is bound to the avatar of the character in the video in the form of a floating layer. When the pressure of the trigger operation is small, less additional information is displayed.

[0131] It can be understood that in the embodiments of the present application, when it comes to data related to user information, etc., when the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.

[0132] The following continues to describe the exemplary structure of the information processing device 455 provided in the embodiments of the present application as software modules. In some embodiments, as Figure 3 shown, the software modules stored in the information processing device 455 in the memory 450 may include: a display module 4551, configured to display image information in a human-computer interaction interface, where the image information includes at least one object; an additional information module 4552, configured to display additional information for describing the object on the image information in response to a first trigger operation on the human-computer interaction interface; where the detail level of the additional information is related to the operation parameters of the first trigger operation.

[0133] In some embodiments, the additional information module 4552 is further configured to: display additional information for describing the object on the image information; where the additional information includes at least one of the following: the association relationship between multiple objects, the attribute information of the object, and the attribute information of the object includes at least one of the following: the role name of the object, the real name of the object.

[0134] In some embodiments, the additional information module 4552 is further configured to: display the object in a marked state in the image information, and display the attribute information of the object at the associated position of the object for the object in the marked state; for two objects with an association relationship, display a connection line for connecting the two objects, and display the association relationship between the two objects at the associated position of the connection line.

[0135] In some embodiments, the additional information module 4552 is further configured to: display an additional information display mode control, where the additional information display mode control is used to switch between an on state and an off state when triggered, the on state indicates that the additional information display mode is turned on, and the off state indicates that the additional information display mode is turned off; and display additional information for describing the object on the image information in response to the additional information display mode control being in the on state and a first trigger operation on the human-computer interaction interface.

[0136] In some embodiments, the additional information module 4552 is further configured to: after displaying additional information for describing the object on the image information, start timing from the display of the additional information; when the timing reaches a first time threshold, display that the additional information display mode control is in the closed state, and at the same time stop displaying all additional information, or stop displaying all additional information in sequence.

[0137] In some embodiments, the image information is a video frame in a video, and the additional information module 4552 is further configured to: display the additional information on the first video frame currently being played in the video; play to the second video frame of the video, where the second video frame is a video frame in the video that comes after the first video frame; when the scene similarity between the second video frame currently being played in the video and the first video frame is less than a scene similarity threshold, stop displaying all additional information at the same time.

[0138] In some embodiments, the additional information module 4552 is further configured to stop displaying all additional information in sequence according to the following order: the order from low to high of the importance level of the additional information; the order from low to high of the importance level of the object to which the additional information belongs.

[0139] In some embodiments, the additional information module 4552 is further configured to perform any one of the following processes: in response to a first trigger operation on a target object in the image information, display additional information of the target object on the image information, where the target object is from the at least one object; in response to a first trigger operation on a target text in the image information, display additional information of the target object on the image information, where the target object is the object specified by the target text and the target object is from the at least one object; in response to a first trigger operation at any position in the image information, display additional information of the target object on the image information, where the target object is the object with the highest importance level among all objects in the image information.

[0140] In some embodiments, the additional information module 4552 is further configured to, before responding to a first trigger operation on the human-computer interaction interface, perform any one of the following processes: in response to a pressing operation on the image information, recognize the pressing operation as a first trigger operation on the human-computer interaction interface; in response to a clicking operation on the image information, recognize the clicking operation as a first trigger operation on the human-computer interaction interface; display an additional information display control, and in response to a control trigger operation on the additional information display control, recognize the control trigger operation as a first trigger operation on the human-computer interaction interface.

[0141] In some embodiments, when the first triggering operation is a non-control triggering operation on the human-computer interaction interface, the operation parameter includes at least one of the following: the contact time of the non-control triggering operation, the contact force of the non-control triggering operation, the contact frequency of the non-control triggering operation; when the first triggering operation is a control triggering operation on the additional information display control, the operation parameter includes: the control parameter of the additional information display control caused by the control triggering operation.

[0142] In some embodiments, the additional information module 4552 is further configured to: in response to meeting the additional information automatic display condition, display additional information for describing the object on the image information; the additional information automatic display condition includes at least one of the following: the number of objects included in the image information exceeds the object number threshold; there is barrage information displayed on the image information, and the barrage information indicates that the audience of the image information has a need to obtain the additional information of the at least one object; the subtitle of the image information includes content related to the at least one object.

[0143] An embodiment of the present application provides a computer program product, which includes computer-executable instructions stored in a computer-readable storage medium. The processor of the electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the information processing method described above in the embodiments of the present application.

[0144] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will be caused to execute the information processing method provided in the embodiments of the present application. For example, Figures 4A - 4C the information processing method shown.

[0145] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.

[0146] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, and may be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0147] By way of example, the computer-executable instructions may or may not correspond to files in a file system, and may be stored as part of a file that holds other programs or data. For example, they may be stored in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program under discussion, or in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code).

[0148] By way of example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed across multiple sites and interconnected via a communication network.

[0149] In summary, by displaying image information including at least one object in a human-computer interaction interface in an embodiment of the present application, where a video including at least one object may be played or an image including at least one object may be displayed, and in response to a first trigger operation on the image information, additional information about the object is displayed on the image information. This is equivalent to directly displaying the additional information about the object through the user's trigger operation, which helps the user efficiently and comprehensively perceive the image information, improves the information representation ability of the image information, and enables the user to obtain the additional information faster and more accurately, thereby improving the human-computer interaction efficiency. Moreover, the level of detail of the additional information is related to the operation parameters of the first trigger operation, which is equivalent to directly controlling the level of detail of the additional information through the operation parameters of the first trigger operation, meeting the user's requirements for the level of detail of the additional information and further improving the human-computer interaction efficiency.

[0150] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.

Claims

1. An information processing method, characterized in that, The method includes: Displaying image information in a human-computer interaction interface, where the image information includes at least one object; In response to a first trigger operation on the human-computer interaction interface, displaying additional information for describing the object on the image information; Wherein, the detail level of the additional information is related to the operation parameters of the first trigger operation.

2. The method according to claim 1, wherein The additional information includes at least one of the following: the association relationship between multiple objects, the attribute information of the object, and the attribute information of the object includes at least one of the following: the role name of the object, the real name of the object.

3. The method according to claim 2, wherein The displaying additional information for describing the object on the image information includes: Displaying the object in a marked state in the image information, and for the object in the marked state, displaying the attribute information of the object at the associated position of the object; For two objects with an association relationship, displaying a connection line for connecting the two objects, and displaying the association relationship between the two objects at the associated position of the connection line.

4. The method according to claim 1, wherein The responding to a first trigger operation on the human-computer interaction interface and displaying additional information for describing the object on the image information includes: Displaying an additional information display mode control, where the additional information display mode control is used to switch between an on state and an off state when triggered, the on state represents the opening of the additional information display mode, and the off state represents the closing of the additional information display mode; In response to the additional information display mode control being in the on state and a first trigger operation on the human-computer interaction interface, displaying additional information for describing the object on the image information.

5. The method according to claim 4, characterized in that, After displaying additional information for describing the object on the image information, the method further includes: Starting timing from the display of the additional information; When the timing reaches a first time threshold, displaying the additional information display mode control in the off state, and at the same time stopping displaying all the additional information, or stopping displaying all the additional information in sequence.

6. The method according to any one of claims 4 or 5, characterized in that The image information is a video frame in a video, The displaying additional information for describing the object on the image information includes: Displaying the additional information on the first video frame currently being played in the video; The simultaneously stopping displaying all the additional information includes: Playing to a second video frame of the video, where the second video frame is a video frame after the first video frame in the video; When the scene similarity between the second video frame currently being played in the video and the first video frame is less than a scene similarity threshold, simultaneously stopping displaying all the additional information.

7. The method according to any one of claims 4 or 5, characterized in that The sequentially stopping displaying all the additional information includes: Stopping displaying all the additional information in sequence according to the following order: The order from low to high of the importance level of the additional information; The order from low to high of the importance level of the object to which the additional information belongs.

8. The method according to claim 1, characterized in that The responding to a first trigger operation on the human-computer interaction interface and displaying additional information for describing the object on the image information includes: Performing any one of the following processes: In response to a first trigger operation on a target object in the image information, additional information of the target object is displayed on the image information, where the target object is from the at least one object; In response to a first trigger operation on a target text in the image information, additional information of a target object is displayed on the image information, where the target object is the object specified by the target text, and the target object is from the at least one object; In response to a first trigger operation on any position in the image information, additional information of a target object is displayed on the image information, where the target object is the object with the highest importance among all the objects in the image information.

9. The method according to claim 1, wherein Before responding to a first trigger operation on the human-machine interface, the method further includes: Performing any one of the following processes: In response to a pressing operation on the image information, identifying the pressing operation as a first trigger operation on the human-machine interface; In response to a clicking operation on the image information, identifying the clicking operation as a first trigger operation on the human-machine interface; Displaying an additional information display control, and in response to a control trigger operation on the additional information display control, identifying the control trigger operation as a first trigger operation on the human-machine interface.

10. The method according to claim 1, wherein when the first trigger operation is a non-control trigger operation on the human-machine interface, the operation parameters include at least one of the following: the contact time of the non-control trigger operation, the contact force of the non-control trigger operation, the contact frequency of the non-control trigger operation; when the first trigger operation is a control trigger operation on an additional information display control, the operation parameters include: the control parameters of the additional information display control caused by the control trigger operation.

11. The method according to claim 1, wherein The method further includes: In response to meeting the additional information automatic display condition, displaying additional information for describing the object on the image information; The additional information automatic display condition includes at least one of the following: the number of objects included in the image information exceeds an object number threshold; there is barrage information displayed on the image information, and the barrage information indicates that the audience of the image information has a need to obtain additional information of the at least one object; the subtitles of the image information include content related to the at least one object.

12. An information processing apparatus, characterized in that, The device includes: a display module, configured to display image information in a human-machine interface, where the image information includes at least one object; an additional information module, configured to, in response to a first trigger operation on the human-machine interface, display additional information for describing the object on the image information; wherein the detail level of the additional information is related to the operation parameters of the first trigger operation.

13. An electronic device, characterized in that, The electronic device includes: a memory, configured to store computer-executable instructions; a processor, configured to implement the information processing method according to any one of claims 1 to 11 when executing the computer-executable instructions stored in the memory.

14. A computer-readable storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed by a processor, the information processing method according to any one of claims 1 to 11 is implemented.

15. A computer program product comprising computer-executable instructions, characterized in that, When the computer-executable instructions are executed by a processor, the information processing method according to any one of claims 1 to 11 is implemented.