Interactive display method and apparatus, electronic device, and readable medium
By collecting users' hand movement trajectories and generating interactive display results using a pre-set database, the problem of monotonous content formats in online teaching is solved, and the fun and interactivity of the display are enhanced.
Patent Information
- Application Number
- CN202111217427.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-19
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2041-10-19
AI Technical Summary
In online teaching scenarios, existing technologies result in user-generated content that is monotonous, fails to vividly and accurately represent the intended message, and lacks engaging elements.
By collecting the movement trajectory of the user's hand, matching and displaying corresponding information using a preset database, the display results are automatically generated, including the object's outline, texture, and related knowledge explanations.
It enables flexible and interactive displays, enhancing the fun and interactivity of the presentations and improving the experience for learners and viewers.
Smart Images

Figure CN116009682B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and more particularly to an interactive display method, apparatus, electronic device, and readable medium. Background Technology
[0002] Many smart devices and applications have interactive display functions, allowing one user to show an object to another in the form of images or videos. For example, after a teacher draws a shape, they can show it to students, who can then learn from and imitate it, greatly facilitating the art teaching process. This process requires the use of image or video processing technology to recognize the content displayed by the user.
[0003] Currently, the recognition of interactive display processes mainly focuses on recognizing and analyzing images drawn by users on screens or drawing boards. In other words, user A needs to draw actual content on the screen or drawing board for user B to view or learn. However, in the context of online teaching through live streaming, this method of displaying on a screen or drawing board has significant limitations. Furthermore, the content drawn by users usually only contains lines, is monotonous, lacks interest, and is difficult to vividly and accurately represent what the user actually wants to express. Summary of the Invention
[0004] This disclosure provides an interactive display method, device, electronic device, and readable medium to achieve convenient and flexible interactive displays and enhance the fun of the display process.
[0005] In a first aspect, embodiments of this disclosure provide an interactive display method, including:
[0006] Collect user's display operation information, including the movement trajectory of the user's hand during the display process;
[0007] The outline of the object is determined based on the motion trajectory;
[0008] Search the preset database for display information that matches the outline;
[0009] The display results are generated based on the corresponding information.
[0010] Secondly, embodiments of this disclosure also provide an interactive display device, including:
[0011] The data acquisition module is used to collect user display operation information, including the movement trajectory of the user's hand during the display process;
[0012] A contour determination module is used to determine the contour of an object based on the motion trajectory.
[0013] The matching module is used to search for display information that matches the outline in a preset database;
[0014] The generation module is used to generate display results based on the corresponding display information.
[0015] Thirdly, embodiments of this disclosure also provide an electronic device, including:
[0016] One or more processors;
[0017] Storage device for storing one or more programs;
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the interactive display method as described in the first aspect.
[0019] Fourthly, embodiments of this disclosure also provide a computer-readable medium having a computer program stored thereon that, when executed by a processor, implements the interactive display method as described in the first aspect.
[0020] This disclosure provides an interactive display method, device, electronic device, and medium. The method collects user display operation information, including the movement trajectory of the user's hand during the display process; determines the outline of an object based on the movement trajectory; searches a preset database for display-corresponding information matching the outline; and generates a display result based on the display-corresponding information. This technical solution can automatically generate a display result based on the user's hand movement trajectory and the display-corresponding information, allowing users greater flexibility and making the display operation more flexible. Furthermore, it provides the necessary display-corresponding information based on the outline displayed by the user, enhancing the fun and interactivity of the interactive display. Attached Figure Description
[0021] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.
[0022] Figure 1 A flowchart illustrating an interactive display method provided in Embodiment 1 of this disclosure;
[0023] Figure 2 This is a schematic diagram illustrating the outline of an object and the corresponding matching information provided in Embodiment 1 of this disclosure;
[0024] Figure 3 This is a flowchart illustrating an interactive display method provided in Embodiment 2 of this disclosure;
[0025] Figure 4 This is a flowchart illustrating an interactive display method provided in Embodiment 3 of this disclosure;
[0026] Figure 5 This is a flowchart illustrating an interactive display method provided in Embodiment 4 of this disclosure;
[0027] Figure 6a This is a schematic diagram of an interactive display based on real-time communication provided in Embodiment 4 of this disclosure;
[0028] Figure 6b This is a schematic diagram of an interactive display based on real-time communication provided in Embodiment 4 of this disclosure;
[0029] Figure 7 This is a schematic diagram of the structure of an interactive display device provided in Embodiment 5 of this disclosure;
[0030] Figure 8 This is a schematic diagram of the structure of an electronic device provided in Embodiment Six of this disclosure. Detailed Implementation
[0031] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0032] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0033] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0034] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0035] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0036] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0037] In the following embodiments, each embodiment provides optional features and examples. The various features described in the embodiments can be combined to form multiple optional solutions. Each numbered embodiment should not be regarded as only one technical solution. Furthermore, unless otherwise specified, the embodiments and features in the embodiments of this disclosure can be combined with each other.
[0038] Example 1
[0039] Figure 1 This is a flowchart illustrating an interactive display method provided in Embodiment 1 of this disclosure. This method is applicable to situations where a display result is automatically generated based on a user-displayed outline to provide to learners or viewers. For example, it can be applied to situations where users are conducting drawing instruction in a live-streaming scenario. This method can be executed by an interactive display device, which can be implemented by software and / or hardware and is generally integrated into an electronic device. In this embodiment, the electronic device includes, but is not limited to, devices such as computers, mobile phones, personal digital assistants, and computers.
[0040] like Figure 1 As shown in Embodiment 1 of this disclosure, an interactive display method includes the following steps:
[0041] S110. Collect user's display operation information, including the movement trajectory of the user's hand during the display process.
[0042] In this embodiment, "user" primarily refers to the presenter, such as a painter, specifically a painting teacher, or any user who conducts painting instruction or other demonstration operations through an electronic device. The demonstration operation mainly refers to the actions performed by the user during the demonstration process to create an object. The object can be understood as the content being displayed, such as what the user has drawn. For example, the demonstration operation could be the user moving their fingers or holding an object (such as chalk or a laser pointer) to form a specific trajectory; it could also be the user using gestures or limbs to draw a specific shape; or it could be the user providing relevant information about the object to the electronic device through voice or command input. The demonstration operation information mainly consists of information collected based on the user's demonstration operations that can be used to determine the object. This information includes at least the movement trajectory of the user's hand during the demonstration. The electronic device primarily determines the object drawn by the user and completes the corresponding teaching process by recognizing the movement trajectory. In addition, the demonstration operation information may include other information to assist the electronic device in more accurately determining the object, such as the trajectory of the user's arm during the demonstration, the shape drawn by the user's hand or limbs, and / or the physical object displayed by the user. For example, if a user makes a heart shape with their hand, and the movement trajectory is basically a heart shape, the object can be more accurately identified as a heart shape based on the shape made; or, if a user is holding an apple, and the movement trajectory is basically the outline of an apple, the object can be more accurately identified as an apple based on the actual object being held.
[0043] The hand can include the user's hand, including the palm and fingers, and can also include the object being held. In this case, the user's hand and the object being held can be regarded as a whole, and this whole can be represented as a point to determine the movement trajectory.
[0044] Specifically, the acquisition process for the user's hand movement trajectory during the demonstration can be described as follows: multiple frames of images of the user during the demonstration are acquired through the image sensor in the electronic device. Each frame of the image contains the user's hand. The hand as a whole is regarded as a point, and the points in each image constitute the hand's movement trajectory in time sequence.
[0045] S120. Determine the outline of the object based on the motion trajectory.
[0046] Specifically, users primarily express the content they draw by tracing outlines during the presentation. For example, if the object is a cat, the user's hand in each frame is considered as a single point, and by connecting the points in each image sequentially, a rough outline of the cat can be obtained.
[0047] S130. Search the preset database for display information that matches the outline.
[0048] The preset database refers to a pre-set database containing various display-related information. This information mainly refers to information associated with the object and available for learners to study. For example, the display-related information includes template objects associated with the object, where the object is predicted based on the user's movement trajectory and corresponds to the user's intended display. Figure 1 For example, the object could be a cat that the user wants to draw; while the template object is a standard-specified template image that can be used for display, such as a preset cat template image. For instance, when the outline of an object is detected to match the outline of a cat, a cat template image can be found in a preset database and displayed as a match. The displayed information may also include texture and color information for the object's outline, as well as text, patterns, or animations to explain relevant knowledge about the object. For example, when the outline of an object matches the outline of a cat, the displayed information could include a template image of a cat with any color or texture, and could also include explanations of the cat's size, breed, food, and habits.
[0049] It should be noted that, based on the outline of the determined object, multiple display information can be matched in the preset database. In order to accurately match one or more display information, the content of the object to be identified in the display operation information or the user's voice commands during the display process can be used to match the display information in a targeted manner.
[0050] S140. Generate display results based on the corresponding information.
[0051] Specifically, based on the corresponding information, display results can be generated. These results mainly refer to the content presented to learners or viewers. Generating display results can involve rendering the outline of an object, such as adding texture or color; adjusting the outline of an object based on a preset database to make it more realistic and aesthetically pleasing; stylizing the outline of an object, such as converting it into a cartoon, oil painting, line drawing, or hand-drawn style; or directly displaying the found text or animation to present the corresponding information.
[0052] Figure 2 This is a schematic diagram illustrating the outline of an object and the corresponding matching information provided in Embodiment 1 of this disclosure. For example... Figure 2 As shown, taking an interactive demonstration in a painting teaching scenario as an example, the left side shows the outline of an object determined by the hand movement trajectory, which can be seen to be a cat; the right side shows a template object matching the outline, in which the cat has color and texture. When generating the demonstration result, this template object can be used directly as the demonstration result, or other corresponding demonstration information can be added on the template object, such as explanations of the object's size, breed, food, and habits.
[0053] This disclosure provides an interactive display method. The method can automatically generate display results based on the movement trajectory of the user's hand and the corresponding display information. Users can freely express themselves, making the display operation more flexible. At the same time, the method can provide the required display information based on the outline displayed by the user, enhancing the fun and interactivity of the interactive display, thereby improving the experience of learners or viewers.
[0054] Example 2
[0055] Figure 3 This is a flowchart illustrating an interactive display method according to Embodiment 2 of this disclosure. Based on the above embodiments, Embodiment 2 specifies the process of collecting user display operation information and searching for display information matching the outline in a preset database.
[0056] In this embodiment, collecting user display operation information includes: acquiring multiple frames of images of the display process using an image acquisition device; performing semantic segmentation on each image to extract the hand region in each image; and generating a motion trajectory based on the hand region in each image. Based on this, by performing semantic segmentation on each acquired image, the hand motion trajectory can be accurately identified, providing a basis for finding the corresponding display information.
[0057] In this embodiment, searching for display correspondence information matching the contour in a preset database includes: determining a template object associated with the contour using a Generative Adversarial Network (GAN); and searching for display correspondence information of the template object in the preset database. Based on this, the GAN can obtain an accurate template object based on the object's contour, thereby finding matching display correspondence information and avoiding discrepancies between the display correspondence information and the object.
[0058] like Figure 3 As shown in Embodiment 2 of this disclosure, an interactive display method includes the following steps:
[0059] S210. Acquire multiple frames of images of the display process using an image acquisition device.
[0060] In this embodiment, the images displayed during the process mainly refer to images containing the user's hands, which can be captured by an image acquisition device (such as a camera, video camera, etc.). The images displayed during the process consist of multiple frames, each frame containing at least the user's hand area, and may also include the user's arm area, objects in the hand, and the background area, etc.
[0061] S220. Perform semantic segmentation on each image to extract the hand region from each image.
[0062] Semantic segmentation classifies each pixel in an image, determines the category of each pixel (such as belonging to the background or foreground object), and thus divides the image into regions.
[0063] In this embodiment, based on semantic segmentation, each pixel in each image of the display process is first classified to determine the category of each point, such as belonging to the hand, arm, handheld object or background, and then the region is divided according to the category to extract the hand region in each image.
[0064] S230. Generate motion trajectories based on the hand regions in each image.
[0065] In this embodiment, a motion trajectory can be generated based on the hand region extracted from each image.
[0066] For example, semantic segmentation methods can be used to learn to segment the hand, arm, object in hand, and background, and then perform multi-target tracking on the bounding rectangles of the masks of different targets. The masks of the hand, object in hand, arm, and background, as well as the queue of bounding rectangles of the object in hand in previous frames and the current frame, can be returned in real time. Based on the queue of bounding rectangles, at least the motion trajectory of the hand can be generated.
[0067] S240. Determine the outline of the object based on the motion trajectory.
[0068] S250. Use GAN to identify template objects associated with the contour.
[0069] The template object can be considered as a complex template image associated with the object's outline (possessing features other than the outline, such as texture and color information), and the appearance of the template image should match the object's outline.
[0070] Specifically, based on the object's outline, a GAN can determine a template object associated with the outline, and then use this determined template object to find the corresponding display information for subsequent template objects. GAN is a deep learning model consisting of two basic neural networks: a generator neural network and a discriminator neural network. Through continuous adversarial training, GAN develops the ability to obtain the desired output based on the input. In this embodiment, GAN, for example, through a pix2pix network, can perform pairwise image transformation; that is, it generates a corresponding template object based on the object's outline and searches a preset database for matching display information.
[0071] One feasible approach to determining the template object is to use a generative adversarial network (GAN) to derive a complex template object from a simple outline formed by hand movements. For example, refer to... Figure 2 A user draws an outline of a cat using their hand. A generative adversarial network can then generate a template object related to the outline of this cat, such as a cat with color and texture.
[0072] S260. Search the preset database for the display information corresponding to the template object.
[0073] Based on the template object determined by GAN, the corresponding display information of the template object can be found in the preset database.
[0074] S270. Generate display results based on the corresponding information.
[0075] The interactive display method in this embodiment can accurately obtain the movement trajectory of the hand by performing semantic segmentation on the collected images, thereby improving the accuracy of contour recognition and providing effective display information and display results. By using generative adversarial networks, the searched display information can better match the contour of the object, avoiding deviations, thereby improving the efficiency and reliability of interactive display and enabling learners to learn the displayed content quickly and accurately.
[0076] As an optional embodiment, after acquiring multiple frames of images from the display process, an optimization is added to perform semantic segmentation on each image to extract non-hand regions from each image; and to correct the motion trajectory based on the non-hand regions in each image.
[0077] The non-hand region can refer to any area in the image other than the hand area, such as the arm area or the area where the object is held. Understandably, due to human or system influences, the motion trajectory generated based on the hand area may have deviations. To ensure the accuracy of trajectory recognition, other regions can be used to correct the motion trajectory.
[0078] Specifically, after acquiring multiple frames of images from the display process, semantic segmentation is first performed on each image to extract non-hand regions. Then, the motion trajectory is corrected based on the non-hand regions in each image. For example, after acquiring multiple frames of images from the display process, semantic segmentation is first performed on each image to extract the arm region. Then, the motion trajectory is corrected based on the arm region in each image. For instance, the hand position can be corrected based on the arm posture in each frame.
[0079] The purpose of adding the correction step in this optional embodiment is to finely adjust the trajectory points on the motion trajectory based on the posture of the non-hand area, in order to improve the accuracy of object contour recognition, based on the motion trajectory generated according to the hand area.
[0080] As an optional embodiment, after extracting the hand region from each image, the following optimization is added: identifying hand pose and / or held object based on the hand region in each image; and determining the category of the object based on the hand pose and / or held object.
[0081] Among them, hand gestures and / or held objects can be understood as those that can roughly reflect the action of the object and / or the object itself. For example, a hand gesture could be a user making a heart shape with their hand, and a held object could be a user holding an apple, etc.
[0082] Specifically, after extracting the hand regions from each image, the hand posture and / or the object being held can be identified based on these regions. Then, the object's category can be determined based on the hand posture and / or the object being held, for example, identifying the object as a cat. Furthermore, outline images of the object can be provided for user reference.
[0083] Based on this optional embodiment, the category of the object can be determined in advance, and then the subsequent step of searching and displaying corresponding information in a preset database based on the object's outline can be performed. The main purpose of this optional embodiment is to narrow down the scope of searching and displaying corresponding information in the preset database, thereby improving the efficiency of the search. Alternatively, the searched and displayed corresponding information can be verified based on the object's category to ensure the correctness of the searched and displayed information.
[0084] For example, a pre-defined database contains cats and dogs, with cats exhibiting different shapes and colors. Based on hand gestures and / or the object being held, the object's category can be roughly determined to be a cat. During this category determination, a rough reference outline image of a cat can be generated first. Then, based on the outline of the cat drawn by the user, it can be confirmed that the object is indeed a cat in a sitting or lying position, thus obtaining the matching display information and generating the display result. Alternatively, display information can be searched based on the object's outline first. Once the matching information is found, the category determined by the hand gestures and / or the object being held can be used to verify its accuracy. For instance, if the template object is a cat in a sitting or lying position, and the category determined by the hand gestures and / or the object being held (the outline reference) is also a cat, the two match, thus verifying the correctness of the display information. This approach avoids errors in display information, such as an object being a cat but the displayed information being a dog.
[0085] In an optional embodiment, a method for generating a display result based on the corresponding display information is further provided, wherein the corresponding display information includes the rendering information of the object, which can be obtained based on the difference between the outline of the object and the template object, and the rendering information includes information that makes the outline of the object have features such as color or texture.
[0086] In this optional embodiment, the outline is rendered based on the rendering information to obtain the display result. The rendering information can be texture and / or coloring information, such as adding patterns and / or colors inside the outline to render the object's outline. When the display information includes the object's rendering information, the outline can be rendered based on the rendering information to obtain the display result, thereby achieving personalized display and enhancing the visual effect of the interactive display process.
[0087] As an optional embodiment, the display result includes a template object and explanatory information about the template object. After generating the display result based on the corresponding display information, the following optimization is added: display the template object in the first area and display the explanatory information in the second area.
[0088] In this embodiment, the first area can be the left side of the screen, and the second area can be the right side of the screen. This optional embodiment does not limit the specific locations of the first and second areas. Generally, the first and second areas do not overlap. The form of the explanatory information is not limited; it can be text, images, animations, or English. Based on this, when the display result includes a template object and its explanatory information, by displaying the template object in the first area and the explanatory information in the second area, a vivid and rich display of corresponding information is achieved, enabling learners or viewers to quickly and accurately learn the displayed content.
[0089] Example 3
[0090] Figure 4 This is a flowchart illustrating an interactive display method provided in Embodiment 3 of this disclosure. Based on the above embodiments, Embodiment 3 specifies the process of searching for and displaying corresponding information in a preset database.
[0091] In this embodiment, before searching for display information matching the outline in a preset database, the method further includes: recognizing keywords in the user's speech stream using an Automatic Speech Recognition (ASR) model; and determining the category of the object based on the keywords. Based on this, the object can be accurately determined by combining the user's hand movement trajectory with the keywords, ensuring the reliability of the interactive display.
[0092] In this embodiment, searching for display correspondence information matching the outline in a preset database includes: filtering display correspondence information of template objects with the same category in the preset database; and searching for display correspondence information matching the outline from the display correspondence information of template objects with the same category. Based on this, determining the object's category first narrows the scope of searching for display correspondence information in the preset database, improving the efficiency and accuracy of the search.
[0093] like Figure 4 As shown in Embodiment 3 of this disclosure, an interactive display method includes the following steps:
[0094] S310. Collect user's display operation information, including the movement trajectory of the user's hand during the display process.
[0095] S320. Determine the outline of the object based on the motion trajectory.
[0096] S330 identifies keywords in the user's speech stream using an Automatic Speech Recognition (ASR) model.
[0097] Keywords primarily refer to words and phrases related to the object spoken by the user, which can be used to help determine the object's category. Keywords can be identified from the user's speech stream. During a single demonstration, there may be one or more objects, and correspondingly, the number of keywords can also be one or more. Keywords are not limited by part of speech; for example, keywords can include nouns, specifically a general term for a type of person or thing, such as "cat," "child," or "flower." If such words are identified in the user's speech stream, they can be used as keywords to provide a basis for determining the object's category. Furthermore, keywords may appear with quantifiers, such as "a cat" or "a flower"; keywords may also appear with verbs related to the demonstration operation, such as "draw," "paint," "describe," "display," or "represent." When the user mentions such verbs, the noun following these verbs may be keywords; keywords can also be short phrases or instructions, such as "draw a cat" or "draw a flower."
[0098] It is understandable that users may engage in voice communication during the presentation, and this voice communication may include keywords. In this embodiment, Automatic Speech Recognition (ASR) technology can be used to convert speech into text in real time for keyword recognition. ASR technology focuses on speech and converts speech signals into corresponding text or commands. This technology can perform text conversion in real time, improving real-time performance, and is particularly suitable for interactive presentations in live streaming scenarios. By recognizing keywords in real time, presentation results can be quickly and automatically generated based on the outline of the object, thereby improving the efficiency and real-time performance of interactive presentations.
[0099] In this implementation, the user's voice during the presentation can be converted into text in real time, and then the converted text can be used to identify keywords. Based on this, according to the keywords and the outline of the object, the corresponding information for the presentation that matches the outline can be found, making the found information more accurate.
[0100] S340. Determine the category of the object based on the keywords.
[0101] S350: Filter the display information corresponding to template objects in the preset database that match the category.
[0102] The above steps determine the object's category. Then, you can filter the preset database to find template objects matching the category, and subsequently search for corresponding information based on outline matching. Objects within the same category may include various colors, breeds, or sizes; therefore, the filtering process can incorporate keywords from voice communication. For example, the object's category might be "cat," but cats within the same category may also include cats with different postures, colors, or textures. Keywords from voice communication can be used for filtering, such as "draw a yellow cat."
[0103] S360. From the display correspondence information of the template object that matches the category, find the display correspondence information that matches the outline.
[0104] For example, matching display information can be selected from the display information of cat-related template objects. Based on this, the scope of searching for display information in the preset database can be narrowed down, improving the efficiency and accuracy of finding display information.
[0105] S370. Generate display results based on the corresponding information.
[0106] The interactive display method in this embodiment, by combining keywords identified by the ASR model with the outline of the object, can make the search for and display of corresponding information more accurate and improve the real-time performance of the interactive display.
[0107] Understandably, determining the category of an object based on keywords in the audio stream can narrow down the scope of searching and displaying corresponding information in a preset database, thereby improving search efficiency. Alternatively, it can also verify the displayed information based on the object's category to ensure the accuracy of the search results.
[0108] The following example illustrates the interactive demonstration process: Taking a painting lesson as an example, the artist starts a live streaming app using an electronic device (such as a computer, tablet, or mobile phone) and says, "Students, please pay attention, I'm going to draw an corn cob." Then, the artist begins to draw, creating an outline of the corn cob by moving their hand. The image capture device collects the teacher's hand movement trajectory in real time, determines the outline of the object as a corn cob based on the movement trajectory, and then searches for the corresponding display information in a pre-set database.
[0109] In addition, the ASR model can also identify keywords in the speech stream containing "a corn", thereby determining the category of the object as "corn". The display correspondence information of template objects that match "corn" in the preset database is filtered. From the display correspondence information that matches "corn", display correspondence information that matches the outline of the object is searched. The final display result includes an image of corn with color or texture. The appearance of the image matches the outline image shape formed by the user's hand movement trajectory.
[0110] As an optional embodiment, the identification of keywords in a user's speech stream using an ASR model can be further optimized as follows: the speech stream is segmented and stored in a buffer; keywords in each segment are identified using an ASR model, and the confidence level of each keyword is determined; the keyword with the highest confidence level is selected as the keyword in the speech stream. Here, confidence level can be understood as the probability that the current keyword is a keyword.
[0111] It should be noted that the user's audio stream during the presentation can contain keywords or non-keywords, such as words unrelated to the presentation operation or words that introduce keywords. In this embodiment, the audio stream is segmented and stored in a buffer, keywords in each segment are identified, and the confidence level of the keywords is determined. Methods for determining keyword confidence level include, for example, whether there are quantifiers before or after the keyword, whether there are specified verbs, or whether there are template objects related to the keyword in a preset library. This embodiment does not limit these methods. For example, "Students, please pay attention" can be stored in the buffer, as this sentence does not contain keywords; while "I will draw a corn cob" can be stored in the buffer, as this sentence contains the keyword "corn" and has a high confidence level. By segmenting and recognizing the audio stream, the processing dimensionality of the audio data can be reduced, interference from irrelevant words can be eliminated, and the efficiency and accuracy of keyword recognition can be improved. Based on this, the keyword with the highest confidence level is taken as the keyword in the audio stream, providing a reliable basis for identifying the object.
[0112] Example 4
[0113] Figure 5This is a flowchart illustrating an interactive display method provided in Embodiment 4 of this disclosure. Based on the above embodiments, Embodiment 4 specifies the implementation of interactive display through the construction of a real-time communication conferencing system.
[0114] In this embodiment, user display operation information is collected, including: acquiring audio and video frames during the display process via a Real-Time Communication (RTC) module; and using callback functions to transmit the video and audio frames to the ASR module and visual effects module respectively via the RTC channel. Based on this, the audio and video signals of the user during the display process can be acquired in real time and transmitted to the corresponding receiving modules, improving the efficiency of interactive displays. Furthermore, by processing audio and video frames separately, keywords are extracted from the audio frames, and the outlines of objects are identified from the video frames; combining these two methods improves the accuracy of the display results.
[0115] In this embodiment, after generating the display result based on the corresponding display information, the method further includes: streaming the display result to the destination device via Aiortc. This ensures that the display result is reliably displayed in real-time on the destination device.
[0116] like Figure 5 As shown in Embodiment 4 of this disclosure, an interactive display method includes the following steps:
[0117] S410: Acquires audio and video frames during the display process via the real-time communication (RTC) module.
[0118] In this embodiment, the displayed operation information includes audio frames and video frames. The audio frames originate from the audio stream signal during the display process, and the video frames can refer to image frames captured during the display process. In this embodiment, the audio and video frames during the display process can be acquired in real time through a Real-Time Communication (RTC) module. The RTC is the foundation for real-time communication, primarily responsible for the real-time transmission of audio and video frames. The RTC module provides encoding and packaging of audio and video frames (i.e., socket transmission), and can also implement the control signaling required for the transmission of audio and video frames, such as publish / subscribe control and bitrate adjustment. On one hand, the RTC module can acquire audio and video frames and send them to the processing module. On the other hand, the processing module sends the processed display results back to the RTC module, which can then publish the display results to the target object.
[0119] S420: Using a callback function, video frames and audio frames are transmitted to the ASR module and visual effects module respectively via the RTC channel.
[0120] The ARS module is used to receive audio frames, which can be used to assist in object recognition. For example, by segmenting the audio frames and recognizing keywords, the category of the object can be determined based on the keywords. The visual effects module is used to receive video frames, which can be used to recognize hand movement trajectories and determine the object outline.
[0121] Specifically, during the initial setup of the RTC module, the RTC channel is added through the RTC front-end to enable the audio and video acquisition function of the electronic device. Then, the audio and video frames of the display process are acquired through the RTC module, and the video and audio frames are transmitted to the ASR module and the visual effects module respectively using callback functions for subsequent motion trajectory processing and object outline determination.
[0122] S430, The results are returned via the visual effects module.
[0123] The display results can be obtained by the hand detection and tracking module and the tracking feedback module. For example, the hand detection and tracking module is used to process video frames and generate the motion trajectory of the hand, and the tracking feedback module is used to determine the outline of the object based on the motion trajectory. The visual effects module can return the display results generated based on the display correspondence information that matches the outline. For example, it can search for the display correspondence information that matches the outline in the preset database and generate the display results based on the display correspondence information.
[0124] S440: Stream the display results from Aiortc to the destination device.
[0125] The target device can refer to the device used to watch the live stream, such as a mobile phone or computer used by students. In this embodiment, the display results obtained in the previous step can be streamed to the target device using Aiortc. Aiortc has a simple and easy-to-implement structure and provides Python bindings, offering channels for exchanging audio, video, and other data.
[0126] Figure 6a This is a schematic diagram illustrating an interactive display based on real-time communication, provided in Embodiment 4 of this disclosure. This embodiment can be implemented using C++ for RTC-based interactive display. Figure 6aAs shown, the left side represents the RTC module, used to acquire Pulse Code Modulation (PCM) audio and video frames during the display process. The RTC module includes an audio frame acquisition submodule and a video frame acquisition submodule. The audio and video frames acquired by these two submodules are transmitted to the ASR module and the visual effects module, respectively. The audio and video frames can be processed by different modules, which can be written as relatively independent static or dynamic libraries. The processing results of the audio frames (i.e., keywords) can assist the visual effects module in determining the corresponding information for display and generating the display results. The combination of these two approaches improves the accuracy of the display. Furthermore, the ASR module and the visual effects module can operate relatively independently and asynchronously to achieve efficient processing.
[0127] Figure 6b This is a schematic diagram illustrating an interactive display based on real-time communication, provided in Embodiment 4 of this disclosure. This embodiment can be implemented using Python for RTC-based interactive display. Figure 6b As shown, the left side is the Aiortc service module, which is used for audio frame tracking and video frame tracking during the display process, as well as pushing the display results returned by the visual effects module to the target device. The Aiortc service module may include an audio frame tracking submodule and a video frame tracking submodule. The audio frames and video frames tracked by the two submodules are transmitted to the ASR module and the visual effects module, respectively. On this basis, the audio frames and video frames can be processed by the same or different processing modules.
[0128] Optionally, the real-time communication system also includes a hand detection and tracking module, which can be used to identify hand areas and movement trajectories, thereby supporting users to freely outline, draw, or create on the screen. The tracking feedback module can be used to draw the outline, find template objects, and display corresponding information back to the visual effects module. Additionally, Lua logic can be used to display the complete result of the drawn outline pattern in color, returned by the algorithm. Furthermore, personalized display of the results can be achieved, such as introducing a virtual whiteboard to display rich corresponding information.
[0129] The interactive display method in this embodiment can acquire the audio and video signals of the user during the display process in real time through the real-time communication RTC module and transmit them to the corresponding receiving module in real time, thereby improving the efficiency of interactive display. The display results are pushed to the target device by Aiortc, ensuring that the display results are reliably displayed on the target device in real time. In addition, personalized display of the display results can be realized to display rich display-related information, thereby improving the fun and interactivity of interactive display.
[0130] Example 5
[0131] Figure 7 This is a schematic diagram of the structure of an interactive display device provided in Embodiment 5 of this disclosure. The device can be implemented by software and / or hardware and is generally integrated into an electronic device.
[0132] like Figure 7 As shown, the device includes:
[0133] The acquisition module 510 is used to acquire user display operation information, including the movement trajectory of the user's hand during the display process;
[0134] Contour determination module 520 is used to determine the contour of the object based on the motion trajectory;
[0135] Matching module 530 is used to search for display corresponding information that matches the outline in a preset database;
[0136] The generation module 540 is used to generate a display result based on the corresponding display information.
[0137] The interactive display device in this embodiment can automatically generate display results based on the user's hand movement trajectory and the corresponding display information. Users can freely express themselves, making the display operation more flexible. At the same time, it can provide the required display information based on the outline displayed by the user, enhancing the fun and interactivity of the interactive display.
[0138] Based on the above, the acquisition module 510 includes:
[0139] The acquisition unit is used to acquire multiple frames of images of the display process through an image acquisition device;
[0140] A hand region extraction unit is used to perform semantic segmentation on each of the images to extract the hand region in each of the images;
[0141] A motion trajectory generation unit is used to generate a motion trajectory based on the hand region in each of the images.
[0142] Based on the above, the matching module 530 includes:
[0143] Template object determination unit, used to determine the template object associated with the contour through GAN;
[0144] The corresponding information search unit is used to search for the corresponding information of the template object in the preset database.
[0145] Based on the above, the acquisition module 510 also includes:
[0146] A non-hand region extraction unit is used to perform semantic segmentation on each of the images to extract the non-hand regions in each of the images;
[0147] A motion trajectory correction unit is used to correct the motion trajectory based on the non-hand areas in each of the images.
[0148] Based on the above, the acquisition module 510 also includes:
[0149] The recognition unit is used to recognize hand posture and / or held object based on the hand region in each of the images;
[0150] A category determination unit is used to determine the category of the object based on the hand posture and / or the handheld object.
[0151] Based on the above, before searching the preset database for display information matching the contour, the device further includes: a voice recognition module, comprising:
[0152] The keyword recognition unit is used to identify keywords in the user's speech stream using an automatic speech recognition (ASR) model.
[0153] A determining unit is used to determine the category of the object based on the keywords.
[0154] Based on the above, the matching module 530 includes:
[0155] A corresponding information filtering unit is used to filter the display corresponding information of template objects that are consistent with the category in the preset database;
[0156] The search unit is used to search for display correspondence information that matches the outline from the display correspondence information of template objects that are consistent with the category.
[0157] Based on the above, the keyword recognition unit is used for:
[0158] The audio stream is segmented and stored in a buffer;
[0159] The keywords in each segment are identified using an ASR model, and the confidence level of the keywords is determined.
[0160] The keyword with the highest confidence level is used as the keyword in the speech stream.
[0161] Based on the above, the acquisition module 510 includes:
[0162] The frame acquisition unit is used to acquire audio and video frames during the display process through the real-time communication (RTC) module.
[0163] The transmission unit is used to transmit the video frame and the audio frame to the ASR module and the visual effects module respectively through the RTC channel using a callback function.
[0164] Based on the above, after generating the display result according to the corresponding display information, the device further includes: a streaming module, comprising:
[0165] The displayed results are then streamed from Aiortc to the target device.
[0166] Based on the above, the displayed corresponding information includes the rendering information of the object;
[0167] Module 540 is generated, including:
[0168] The outline is rendered based on the rendering information to obtain the display result.
[0169] Based on the above, the display result includes a template object and explanatory information about the template object;
[0170] After generating the display result based on the corresponding display information, the device further includes: a display module, comprising:
[0171] The template object is displayed in the first area, and the explanatory information is displayed in the second area.
[0172] The interactive display device described above can execute the interactive display method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.
[0173] Example 6
[0174] Figure 8 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of this disclosure. Figure 8 A schematic diagram of the structure of an electronic device 600 suitable for implementing embodiments of the present disclosure is shown. The electronic device 600 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 8 The illustrated electronic device 600 is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.
[0175] like Figure 8As shown, electronic device 600 may include one or more processing devices (e.g., central processing unit, graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. One or more processing devices 601 implement the methods provided in this disclosure. Various programs and data required for the operation of electronic device 600 are also stored in RAM 603. Processing devices 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0176] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 608 including, for example, magnetic tape, hard disk, etc., for storing one or more programs; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0177] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0178] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0179] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as Hypertext Transfer Protocol (HTTP), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0180] The aforementioned computer-readable medium may be included in the aforementioned electronic device 600; or it may exist independently and not assembled into the electronic device 600.
[0181] The aforementioned computer-readable medium stores one or more computer programs that, when executed by a processing device, implement the following method: the computer-readable medium carries one or more programs that, when executed by an electronic device, cause the electronic device 600 to: write computer program code for performing the operations of this disclosure in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on a user's computer, partially on a user's computer, as a standalone software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0182] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. Each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0183] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the functionality of the module itself.
[0184] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programming Logic Device (CPLD), and so on.
[0185] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0186] According to one or more embodiments of this disclosure, Example 1 provides an interactive display method, including:
[0187] Collect user's display operation information, including the movement trajectory of the user's hand during the display process;
[0188] The outline of the object is determined based on the motion trajectory;
[0189] Search the preset database for display information that matches the outline;
[0190] The display results are generated based on the corresponding information.
[0191] According to one or more embodiments of this disclosure, Example 2 describes the method described in Example 1.
[0192] The information collected from the user's display operations includes:
[0193] Multiple frames of images of the display process are captured using image acquisition equipment;
[0194] Semantic segmentation is performed on each of the images to extract the hand region from each of the images;
[0195] A motion trajectory is generated based on the hand region in each of the images.
[0196] According to one or more embodiments of this disclosure, Example 3 describes the method described in Example 1.
[0197] Search the preset database for display information that matches the outline, including:
[0198] The template object associated with the contour is determined by a generative adversarial network (GAN);
[0199] Search the preset database for the display information corresponding to the template object.
[0200] According to one or more embodiments of this disclosure, Example 4 describes the method described in Example 2.
[0201] The method further includes:
[0202] Semantic segmentation is performed on each of the images to extract non-hand regions from each of the images;
[0203] The motion trajectory is corrected based on the non-hand areas in each of the images.
[0204] According to one or more embodiments of this disclosure, Example 5 describes the method described in Example 2.
[0205] The method further includes:
[0206] Identify hand posture and / or held object based on the hand region in each of the images;
[0207] The category of the object is determined based on the hand posture and / or the object being held.
[0208] According to one or more embodiments of this disclosure, Example 6 describes the method described in Example 1.
[0209] Before searching the preset database for display information that matches the outline, the process also includes:
[0210] The user's speech stream is identified using an Automatic Speech Recognition (ASR) model;
[0211] The category of the object is determined based on the keywords.
[0212] According to one or more embodiments of this disclosure, Example 7 describes the method according to Example 5 or 6.
[0213] Search the preset database for display information that matches the outline, including:
[0214] Filter the display information corresponding to template objects that match the category in the preset database;
[0215] From the display correspondence information of template objects that are consistent with the category, find the display correspondence information that matches the outline.
[0216] According to one or more embodiments of this disclosure, Example 8 describes the method according to Example 6.
[0217] The ASR model identifies keywords in the user's speech stream, including:
[0218] The audio stream is segmented and stored in a buffer;
[0219] The keywords in each segment are identified using an ASR model, and the confidence level of the keywords is determined.
[0220] The keyword with the highest confidence level is used as the keyword of the speech stream.
[0221] According to one or more embodiments of this disclosure, Example 9 describes the method described in Example 1.
[0222] The information collected from the user's display operations includes:
[0223] The audio and video frames of the display process are acquired through the real-time communication (RTC) module.
[0224] Using a callback function, the video frame and the audio frame are transmitted to the ASR module and the visual effects module respectively via the RTC channel.
[0225] According to one or more embodiments of this disclosure, Example 10 describes the method according to Example 9.
[0226] After generating the display result based on the corresponding display information, the process also includes:
[0227] The displayed results are then streamed from Aiortc to the target device.
[0228] According to one or more embodiments of this disclosure, Example 11 describes the method described in Example 1.
[0229] The displayed information includes the rendering information of the object;
[0230] The step of generating a display result based on the corresponding display information includes:
[0231] The outline is rendered based on the rendering information to obtain the display result.
[0232] According to one or more embodiments of this disclosure, Example 12 describes the method described in Example 1.
[0233] The display results include a template object and explanatory information about the template object;
[0234] After generating the display result based on the corresponding display information, the method further includes:
[0235] The template object is displayed in the first area, and the explanatory information is displayed in the second area.
[0236] According to one or more embodiments of this disclosure, Example 13 provides an interactive display device, including:
[0237] The data acquisition module is used to collect user display operation information, including the movement trajectory of the user's hand during the display process;
[0238] A contour determination module is used to determine the contour of an object based on the motion trajectory.
[0239] The matching module is used to search for display information that matches the outline in a preset database;
[0240] The generation module is used to generate display results based on the corresponding display information.
[0241] According to one or more embodiments of this disclosure, Example 14 provides an electronic device comprising:
[0242] One or more processing devices;
[0243] Storage device for storing one or more programs;
[0244] When the one or more programs are executed by the one or more processing devices, the one or more processing devices implement the interactive display method as described in any of Examples 1-12.
[0245] According to one or more embodiments of the present disclosure, Example 15 provides a computer-readable medium having a computer program stored thereon that, when executed by a processing device, implements the interactive display method as described in any of Examples 1-12.
[0246] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0247] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0248] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. An interactive display method, characterized in that, include: Collect user's display operation information, including the movement trajectory of the user's hand during the display process; The outline of the object is determined based on the motion trajectory; The user's speech stream is identified using an Automatic Speech Recognition (ASR) model; The category of the object is determined based on the keywords; Search the preset database for display information that matches the outline; The display results are generated based on the corresponding display information. The displayed corresponding information includes the rendering information of the object, which is obtained based on the difference between the outline and the template object associated with the outline, including information that makes the outline of the object have color or texture features; The step of generating a display result based on the corresponding display information includes: The outline is rendered based on the rendering information to obtain the display result; Search the preset database for display information that matches the outline, including: Filter the display information corresponding to template objects that match the category in the preset database; From the display correspondence information of template objects that are consistent with the category, find the display correspondence information that matches the outline.
2. The method according to claim 1, characterized in that, The information collected from the user's display operations includes: Multiple frames of images of the display process are captured using image acquisition equipment; Semantic segmentation is performed on each of the images to extract the hand region from each of the images; A motion trajectory is generated based on the hand region in each of the images.
3. The method according to claim 1, characterized in that, Search the preset database for display information that matches the outline, including: The template object is determined by a Generative Adversarial Network (GAN); Search the preset database for the display information corresponding to the template object.
4. The method according to claim 2, characterized in that, The method further includes: Semantic segmentation is performed on each of the images to extract non-hand regions from each of the images; The motion trajectory is corrected based on the non-hand areas in each of the images.
5. The method according to claim 2, characterized in that, The method further includes: Identify hand posture and / or held object based on the hand region in each of the images; The category of the object is determined based on the hand posture and / or the object being held.
6. The method according to claim 1, characterized in that, The ASR model identifies keywords in the user's speech stream, including: The audio stream is segmented and stored in a buffer; The keywords in each segment are identified using an ASR model, and the confidence level of the keywords is determined. The keyword with the highest confidence level is used as the keyword in the speech stream.
7. The method according to claim 1, characterized in that, The information collected from the user's display operations includes: The audio and video frames of the display process are acquired through the real-time communication (RTC) module. Using a callback function, the video frame and the audio frame are transmitted to the ASR module and the visual effects module respectively via the RTC channel.
8. The method according to claim 7, characterized in that, After generating the display result based on the corresponding display information, the process also includes: The displayed results are then streamed from Aiortc to the target device.
9. The method according to claim 1, characterized in that, The display results include a template object and explanatory information about the template object; After generating the display result based on the corresponding display information, the method further includes: The template object is displayed in the first area, and the explanatory information is displayed in the second area.
10. An interactive display device, characterized in that, include: The data acquisition module is used to collect user display operation information, including the movement trajectory of the user's hand during the display process; A contour determination module is used to determine the contour of an object based on the motion trajectory. The matching module is used to search for display information that matches the outline in a preset database; The generation module is used to generate display results based on the corresponding display information; The displayed corresponding information includes the rendering information of the object, which is obtained based on the difference between the outline and the template object associated with the outline, including information that makes the outline of the object have color or texture features; The generation module is specifically used for: The outline is rendered based on the rendering information to obtain the display result; The interactive display device also includes: a voice recognition module. The speech recognition module includes: The keyword recognition unit is used to identify keywords in the user's speech stream using an automatic speech recognition (ASR) model. A determining unit, configured to determine the category of the object based on the keywords; The matching module includes: A corresponding information filtering unit is used to filter the display corresponding information of template objects that are consistent with the category in the preset database; The search unit is used to search for display correspondence information that matches the outline from the display correspondence information of template objects that are consistent with the category.
11. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the interactive display method as described in any one of claims 1-9.
12. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the interactive display method as described in any one of claims 1-9.
Citation Information
Patent Citations
Digital virtual-real interaction system and digital virtual-real interaction method
CN103823554A
Animal imitation processing method and device based on gesture recognition, equipment and medium
CN112766231A