Image tagging method in conjunction with sound signals, terminal device and server
By embedding audio signals into video conferencing, the terminal device and server can mark target areas in the image, solving the problem of ineffective marking in the prior art and improving the convenience of video conferencing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ACER INC
- Filing Date
- 2022-09-26
- Publication Date
- 2026-07-24
AI Technical Summary
In video conferencing, other participants cannot add special prompts to the content projected by the speaker, making it difficult for existing technologies to effectively indicate or mark target areas when it is necessary to explain specific parts.
By combining sound signals with image tagging, the terminal device and the server embed the target sound signal into the speech signal to generate a combined sound signal, and indicate the target area in the image. The terminal device and the server process and transmit the image signal respectively to achieve tagging.
It enhances the convenience of video conferencing, enabling other participants to indicate target areas in images via audio signals, thus improving the experience of multi-person meetings.
Smart Images

Figure CN117812215B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a signal processing technology, and more particularly, to an image tagging method, terminal device, and server that incorporates sound signals. Background Technology
[0002] Remote conferencing allows multiple people in different locations or spaces to converse, and the related equipment, protocols, and applications are quite mature. It's worth noting that during a video conference, the presenter's computer can share / project its screen so other participants can view the desktop, documents, or specific applications. However, according to the settings provided by current video conferencing software, participants can only see the content projected by the presenter; they cannot add specific prompts or comments to the presenter's projection. When other users want to explain a specific part of the projected content, they still need to painstakingly explain that specific part. Summary of the Invention
[0003] This invention relates to an image marking method, terminal device, and server that incorporates sound signals, thereby improving convenience by carrying instructions for marking on images via sound signals.
[0004] According to embodiments of the present invention, an image tagging method combining sound signals includes (but is not limited to) the following steps: displaying a first image; detecting a selection instruction; embedding a target sound signal into a speech signal to generate a combined sound signal; and transmitting the combined sound signal. The selection instruction corresponds to a target region in the first image, and the selection instruction is generated by selecting the target region through an input operation. The target sound signal corresponds to the target region of the selection instruction, and the speech signal is obtained through audio recording.
[0005] According to an embodiment of the present invention, the terminal device includes (but is not limited to) a display, a communication transceiver, an input device, a memory, and a processor. The memory stores program code. The processor is coupled to the display, the communication transceiver, the input device, and the memory. The processor is configured to load the program code to perform the following steps: displaying a first image; detecting a selection instruction; embedding a target sound signal into a speech signal to generate a combined sound signal; and transmitting the combined sound signal. The selection instruction corresponds to a target area in the first image, and the selection instruction is generated by selecting the target area through an input operation. The target sound signal corresponds to the target area of the selection instruction, and the speech signal is obtained through audio recording.
[0006] According to an embodiment of the present invention, the server includes (but is not limited to) a communication transceiver, a memory, and a processor. The memory stores program code. The processor is coupled to the communication transceiver and the memory. The processor is configured to load the program code to perform the following steps: receiving a combined audio signal; distinguishing the combined audio signal into a speech signal and a target audio signal; determining a target region corresponding to the target audio signal; generating a marker in the target region of a second image to generate a first image signal; and transmitting the first image signal. The speech signal is obtained through audio reception. The first image signal includes a second image with the marker.
[0007] Based on the above, according to the image tagging method, terminal device, and server combining audio signals according to embodiments of the present invention, the terminal device can embed a target audio signal corresponding to a target area in an image into an audio signal, and the server can add a mark to the target area in the image according to the target audio signal. Thus, under the settings of the video software, by embedding image tags with audio signals, convenience is improved, thereby enhancing the video conferencing experience. Attached Figure Description
[0008] The accompanying drawings are included to further illustrate the invention, and are incorporated in and constitute a part of this specification. The drawings illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.
[0009] Figure 1 This is a component block diagram of a system according to an embodiment of the present invention;
[0010] Figure 2 This is a component block diagram of a terminal device according to an embodiment of the present invention;
[0011] Figure 3 This is a component block diagram of a server according to an embodiment of the present invention;
[0012] Figure 4 This is a flowchart of an image tagging method for a terminal device incorporating audio signals according to an embodiment of the present invention;
[0013] Figure 5 This is a schematic diagram of a user interface for video software according to an embodiment of the present invention;
[0014] Figure 6 This is a schematic diagram of region segmentation according to an embodiment of the present invention;
[0015] Figure 7 This is a flowchart illustrating the generation of instructions for triggering operations according to an embodiment of the present invention;
[0016] Figure 8 This is a flowchart of matching, filtering, and embedding according to an embodiment of the present invention;
[0017] Figure 9 This is a flowchart illustrating the generation of a cancellation operation instruction according to an embodiment of the present invention;
[0018] Figure 10 This is a flowchart of an image tagging method for a server that incorporates audio signals, according to an embodiment of the present invention.
[0019] Figure 11 This is a flowchart of filtering, matching, and labeling according to an embodiment of the present invention;
[0020] Figure 12 This is a schematic diagram of mark generation according to an embodiment of the present invention;
[0021] Figure 13 This is a schematic diagram of the cancellation of a mark according to an embodiment of the present invention.
[0022] Explanation of icon numbers
[0023] 1: System;
[0024] 10: Terminal device;
[0025] 11: Monitor;
[0026] 12: Communication transceiver;
[0027] 13: Input devices;
[0028] 14: Memory;
[0029] 15: Processor;
[0030] 16: Microphone;
[0031] 30: Server;
[0032] 33: Communication transceiver;
[0033] 34: Memory;
[0034] 35: Processor;
[0035] 50: Network;
[0036] S410~S440, S710~S730, S810~S840, S910~S930, S101~S105, S111~S116: Steps;
[0037] SC: Share screen;
[0038] UI: User Interface;
[0039] C1, C2: Vernier;
[0040] A: Region;
[0041] C A :Select command;
[0042] Target sound signal;
[0043] S mic S tx :Original sound signal;
[0044] Voice signal;
[0045] x1, x2, ..., x N Combined sound signals;
[0046] x: Synthesized speech signal;
[0047] y: First image signal;
[0048] M1, M2: Markers. Detailed Implementation
[0049] Reference will now be made in detail to exemplary embodiments of the invention, examples of which are illustrated in the accompanying drawings. Wherever possible, the same component reference numerals are used in the drawings and description to denote the same or similar parts.
[0050] Figure 1 This is a component block diagram of system 1 according to an embodiment of the present invention. Please refer to... Figure 1 System 1 includes (but is not limited to) one or more terminal devices 10 and servers 30.
[0051] Terminal device 10 may be a mobile phone, VoIP phone, tablet computer, desktop computer, laptop computer, smart assistant, or in-vehicle system.
[0052] Figure 2 This is a component block diagram of a terminal device according to an embodiment of the present invention. Please refer to... Figure 2 The terminal device 10 includes (but is not limited to) a display 11, a communication transceiver 12, an input device 13, a memory 14, and a processor 15.
[0053] Display 11 may be a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic light-emitting diode (OLED), a quantum dot display, or other types of display. In one embodiment, display 11 is used to display images, such as user interfaces, documents, pictures, or videos.
[0054] The communication transceiver 12 can be a communication transceiver supporting technologies such as fourth-generation (4G), fifth-generation (5G) or other generations of mobile communications, Wi-Fi, Bluetooth, infrared, radio frequency identification (RFID), Ethernet, fiber optic networks, serial communication interfaces (e.g., RS-232), or Universal Serial Bus (USB), Thunderbolt, or other communication transmission interfaces. In this embodiment of the invention, the communication transceiver 12 is used to transmit or receive data with other electronic devices (e.g., server 30 or other terminal devices 10) via network 50 (e.g., wired network, wireless network, or private network).
[0055] Input device 13 can be a mouse, keyboard, touch panel, trackball, button, or switch. In one embodiment, input device 13 is used to receive input operations (e.g., swipe, press, touch, or toggle operations) and generate corresponding instructions accordingly. It should be noted that input operations on multiple components of input device 13 may generate different instructions. For example, pressing the left mouse button generates a selection instruction. As another example, clicking the right mouse button twice generates a cancel instruction. The content and function of the instructions will be described in subsequent embodiments.
[0056] The memory 14 can be any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid-state drive (SSD), or similar component. In one embodiment, the memory 14 is used to store program code, software modules, configuration settings, data (e.g., images, instructions, sound signals, etc.) or files, and embodiments thereof will be described in detail later.
[0057] Processor 15 is coupled to display 11, transceiver 12, input device 13, and memory 14. Processor 15 may be a central processing unit (CPU), graphics processing unit (GPU), or other programmable general-purpose or special-purpose microprocessor, digital signal processor (DSP), programmable controller, field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), neural network accelerator, or other similar components or combinations thereof. In one embodiment, processor 15 is used to execute all or part of the operations of terminal device 10, and can load and execute program code, software modules, files, and data stored in memory 14. In some embodiments, the functionality of processor 15 may be implemented via software or a chip.
[0058] In one embodiment, the terminal device 10 further includes a microphone 16. The microphone 16 may be a dynamic, condenser, or electret condenser type microphone. The microphone 16 may also be a combination of other electronic components, analog-to-digital converters, filters, and audio processors that receive sound waves (e.g., human voices, ambient sounds, machine noises, etc.) and convert them into sound signals. In one embodiment, the microphone 16 is used to pick up / record the speaker's voice to obtain a speech signal. In some embodiments, this speech signal may include the speaker's voice, sounds emitted by a speaker, and / or other ambient sounds.
[0059] Figure 3 This is a block diagram of the components of a server 30 according to an embodiment of the present invention. Please refer to... Figure 3 The server 30 includes (but is not limited to) a communication transceiver 33, a memory 34, and a processor 35.
[0060] The implementation and functions of the communication transceiver 33, memory 34, and processor 35 can be referred to the descriptions of the communication transceiver 13, memory 14, and processor 15, respectively, and will not be repeated here. In one embodiment, the processor 35 is used to execute all or part of the operations of the server 30, and can load and execute the program code, software modules, files, and data stored in the memory 34.
[0061] The methods described in the embodiments of the present invention will be explained below in conjunction with the various devices, components and modules in System 1. The various processes of this method may be adjusted according to the implementation situation, and are not limited thereto.
[0062] Figure 4 This is a flowchart of an image tagging method for a terminal device 10 incorporating audio signals according to an embodiment of the present invention. Please refer to... Figure 4 The processor 15 displays the first image via the display 11 (step S410). In one embodiment, the first image may be a user interface of video software (e.g., Zoom, Webex, Teams, or Meet). For example, Figure 5 This is a schematic diagram of the user interface (UI) of video software according to an embodiment of the present invention. Please refer to... Figure 5 Depending on different design requirements, the user interface (UI) can display participant illustrations or real-time images, as well as a shared screen (SC). The content of the shared screen (SC) can be, for example, a slide, a document, a video, or an image. In another embodiment, the first image may also be a screenshot of a user interface of other types of software, a streaming image, a video, an image, or a document.
[0063] Processor 15 detects a selection instruction (step S420). Specifically, the selection instruction corresponds to a target region in the first image, and the selection instruction is generated by selecting a target region through an input operation received by input device 13. In other words, the first image includes one or more regions, and the input operation is used to select a target region among those regions of the first image.
[0064] For example, Figure 6 This is a schematic diagram of region segmentation according to an embodiment of the present invention. Please refer to... Figure 6 , Figure 5 The sharing screen SC in the user interface (UI) is divided into multiple regions A. Adjacent regions A may not overlap or may partially overlap. Regions A are labeled "1", "2", ..., "16" to represent their identifiers. Assume that the cursor C1 of another user (e.g., a secondary terminal device 10) is located in region A labeled "1", while the cursor C2 of the speaker (e.g., the primary terminal device 10) is located in region A labeled "6". These regions A include the target region. When the secondary terminal device 10 also receives a double-click input, region A labeled "1" becomes the target region. It should be noted that... Figure 6 The segmentation method and naming conventions of the illustrated regions are merely illustrative examples, and users can modify the segmentation method and identifiers according to their actual needs. For example, the identifiers can be coordinates in a two-dimensional coordinate system.
[0065] Figure 7This is a flowchart illustrating the generation of instructions for a trigger operation according to an embodiment of the present invention. Please refer to... Figure 7 The processor 15 can compare the input operation with the trigger operation to generate a first comparison result (step S710). The trigger operation can be one or more preset operations. For example, clicking the left mouse button once, touching a certain area in the first image, or a specific button. Another example is that the 16 keys on the keyboard correspond to... Figure 6 The system has 16 regions, and the trigger operation can be the pressing of any one of these 16 buttons. However, the definition of the trigger operation may vary, and users can change it according to their actual needs; this embodiment of the invention does not impose any limitations on this. The processor 15 determines whether the input operation is a preset trigger operation. Therefore, the first comparison result includes whether the input operation matches / is the same as the trigger operation, and whether the input operation does not match / is different from the trigger operation.
[0066] In response to an input operation that matches or is identical to a trigger operation, processor 15 can determine the target area selected by the input operation (step S720). For example, processor 15 determines the area where the cursor is located, or the area corresponding to a specific component of input device 13 (e.g., a key, button, or sensing component). Figure 6 For example, other users' cursors C1 are located in region A of the identifier "1", so region A is the target region.
[0067] Processor 15 can generate a selection instruction based on the target region (step S730). Since the location of the target region in the first image has been confirmed, the selection instruction is an instruction regarding the selection of the target region, and the selection instruction is detected accordingly. In response to an input operation that does not conform to / is different from the trigger operation, the target region is disabled / stopped / determined and / or a selection instruction is generated.
[0068] Please refer to Figure 4 The processor 15 embeds the target sound signal into the speech signal to generate a combined sound signal (step S430). Specifically, the target sound signal corresponds to the target area of the selection instruction, and the speech signal is obtained through radio reception.
[0069] Figure 8 This is a flowchart illustrating matching, filtering, and embedding according to an embodiment of the present invention. Please refer to... Figure 8 Processor 15 can select instruction C A Determine the target sound signal that matches the identifier of the target region from one or more sample sound signals. (Step S810). One or more regions each correspond to one or more identifiers. For example, Figure 6The 16 regions A shown correspond to identifiers "1" through "16". One or more identifiers also correspond to one or more sample sound signals. The sample sound signal can be any custom sound signal, such as a sound signal with a specific frequency band, encoding, amplitude, waveform, or melody. Different identifiers correspond to different sample sound signals. That is, there is a one-to-one correspondence between regions and sample sound signals. However, in other embodiments, there may also be a many-to-one or one-to-many correspondence between regions and sample sound signals. Figure 6 For example, the processor 15 may select the sample sound signal with identifier "1" as the target sound signal. In some embodiments, the processor 15 may directly find the target sound signal matching the target region based on the correspondence between the region and the sample sound signal.
[0070] On the other hand, the processor 15 can pick up sound via the microphone 16 or receive raw sound signals S from other recording devices. mic That is, the original sound signal S mic It is the sound signal generated by recording / receiving sound from a sound source (e.g., a user, animal, or environment). Processor 15 can process the original sound signal S mic Echo cancellation, noise suppression, power gain, and / or audio signal processing are performed (step S820, optionally) to generate the original audio signal S. tx Processor 15 can process the original sound signal S tx The speech signal is generated by passing through a filter (step S830). This filter is used to filter out sound signals outside the first frequency band, and speech signals It belongs to the first frequency band. For example, the first frequency band is frequencies below 5kHz or frequencies between 2kHz and 5kHz. The target sound signal... It belongs to the second frequency band, which is higher than the first frequency band. For example, the second frequency band is the frequency between 5kHz and 8kHz or the frequency above 6kHz.
[0071] Next, the processor 15 can convert the target sound signal Embedded voice signal (Step S840). For example, the processor 15 can directly superimpose the target sound signal in the time domain or frequency domain. and voice signals This outputs a combined sound signal x1.
[0072] Please refer to Figure 4The processor 15 transmits the combined audio signal via the network 50 through the transceiver 12 (step S440). Specifically, compared to the prior art which directly transmits the voice signal, the target voice signal in the combined audio signal of this embodiment can correspond to a target area in the first image, thereby indicating that the target area is selected or needs to be focused / emphasized / marked. In addition, if no selection instruction is detected, the terminal device 10 may also directly transmit the voice signal.
[0073] Image tagging can be processed by server 30, as detailed in subsequent embodiments. Next, processor 15 can receive image signals from server 30 or other devices. Processor 15 can display a second image from the image signals on display 11. The second image is a shared frame (e.g., a video image, streaming image, movie, picture, or file frame). The target area in this second image is marked. The mark can be any pattern, shape, color, symbol, transparency, and / or texture. For example, stars, hearts, or squares. A detailed description of the image signals will also be provided in subsequent embodiments.
[0074] In addition to indicating the target area to be selected or to be focused on / emphasized / marked, you can further deselect / focus / emphasize / mark it. Figure 9 This is a flowchart illustrating the generation of a cancellation operation instruction according to an embodiment of the present invention. Please refer to... Figure 9 The processor 15 can compare the input operation with the cancellation operation to generate a second comparison result (step S910). Similarly, the cancellation operation can be one or more preset operations. For example, clicking the right mouse button once, touching a target area in a marked second image, or a specific button. Or, for example, the 16 keys of the keyboard correspond to... Figure 6 The system has 16 regions, and a cancellation operation can be achieved by pressing any one of these 16 buttons twice. However, the definition of a cancellation operation can vary, and users can modify it according to their actual needs; this embodiment of the invention does not impose any limitations on this. The processor 15 determines whether the input operation is a preset cancellation operation. Therefore, the second comparison result includes whether the input operation matches / is the same as the trigger operation, and whether the input operation does not match / is different from the trigger operation.
[0075] In response to an input operation that matches or cancels the input operation, processor 15 can determine the target area selected by the input operation (step S920). For example, processor 15 determines the area where the cursor is located, or the area corresponding to a specific component of input device 13 (e.g., a key, button, or sensing component). Figure 6 For example, other users' cursors C1 are located in region A of the identifier "1", so region A is the target region.
[0076] Processor 15 can generate a selection instruction based on the target region (step S930). Since the location of the target region in the second image has been confirmed, the selection instruction is generated to select the target region, and the selection instruction is detected accordingly. Furthermore, with... Figure 7 The difference in the implementation of the triggering operation is that the selection instruction is accompanied by a cancellation instruction, which is used to remove the marking of the target region in the (marked) second image. In response to an input operation that does not conform to / is different from the cancellation operation, the target region is disabled / stopped / undefined and / or a selection instruction is generated.
[0077] It should be noted that the embodiments of the present invention are not limited to the (secondary) terminal device 10 of other users who are not sharing the screen transmitting the combined audio signal of the marker indication. The (primary) terminal device 10 of the speaker who is sharing the screen may also transmit the combined audio signal of the marker indication as needed.
[0078] Figure 10 This is a flowchart of an image tagging method for a server 30 incorporating audio signals according to an embodiment of the present invention. Please refer to... Figure 10 The processor 35 receives the combined audio signal via the network 50 through the communication transceiver 33 (step S101). The combined audio signal is the audio signal transmitted by the aforementioned terminal device 10 via the network 50.
[0079] Processor 35 separates the combined sound signal into a speech signal and a target sound signal (step S102). Figure 4 As can be seen from the embodiments, the combined sound signal is generated by embedding the target sound signal into the speech signal. Therefore, the processor 35 separates the speech signal and the target sound signal from the combined sound signal to provide different subsequent processing.
[0080] Figure 11 This is a flowchart of filtering, matching, and labeling according to an embodiment of the present invention. Please refer to... Figure 11 The combined sound signals x1, x2, ..., x received in step S111 N (N is a positive integer) represent the audio signals transmitted by different terminal devices 10 via network 50. Taking the combined audio signal x1 as an example, the processor 35 can pass the combined audio signal x1 through the first filter (step S112) to generate a speech signal. (For example, Figure 8 voice signal Similarly, the first filter is used to filter out sound signals outside the first frequency band, and the speech signal belongs to the first frequency band. For example, the first frequency band is a frequency below 5 kHz or a frequency between 2 kHz and 5 kHz. In one embodiment, the processor 35 can filter different combinations of sound signals x1, x2, ..., x... NThe separated / distinguished speech signal Perform processes such as synthesis, superposition, echo cancellation, noise suppression and / or other audio signal processing (step S113) to generate a synthesized speech signal x.
[0081] On the other hand, the target sound signal The second frequency band is higher than the first frequency band. For example, the second frequency band is a frequency between 5kHz and 8kHz or a frequency above 6kHz. Taking the combined sound signal x1 as an example, the processor 35 can pass the combined sound signal x1 through the second filter (step S114) to generate the target sound signal. (For example, Figure 8 voice signal The second filter is used to filter out audio signals outside the second frequency band, so the output of the second filter can retain the target audio signal.
[0082] Please refer to Figure 10 The processor 35 determines the target region corresponding to the target sound signal (step S103). Specifically, as follows: Figure 4 As described in the embodiments, each sample sound signal corresponds to one or more regions in the first image. Please refer to... Figure 11 The processor 35 can determine that a first sample sound signal among one or more sample sound signals matches the target sound signal, and determine the target region based on the identifier of the first sample sound signal (step S115). For example, the processor 35 can use cross-correlation or other techniques to compare sound signals to determine the correlation between the target sound signal and any sample sound signal, and use the sample sound signal with the highest correlation or similarity as the first sample sound signal. That is, the first sample sound signal with the highest correlation / similarity matches the target sound signal.
[0083] Furthermore, one or more regions each correspond to one or more identifiers. For example, Figure 6 The 16 regions A shown correspond to identifiers “1” to “16”, respectively. One or more identifiers also correspond to one or more sample audio signals. When the second image is the user interface of video software (e.g., Figure 5 The user interface (UI) shown can be divided into multiple areas (e.g., the sharing screen within the user interface). Figure 6The shared screen SC shown is divided into 16 regions. The second image is the image in the image signal received by the terminal device 10 as described in the aforementioned embodiment of the terminal device 10. That is, the second image is the image in the image signal that will be generated and / or transmitted by the server 30, and the second image is also the screen to be shared. If the target audio signal is identified as one of one or more sample audio signals, the processor 35 can also determine which of one or more regions the target region is based on the correspondence between the sample audio signals and the regions. For example, Figure 12 This is a schematic diagram illustrating the generation of a mark according to an embodiment of the present invention. Please refer to... Figure 12 Assume the target region is the region identified by the identifier "1".
[0084] Please refer to Figure 10 The processor 35 generates markers in the target area of the second image to generate a first image signal (step S104). The second image is an image provided to the terminal device 10 participating in the same video conference. The processor 35 can draw, add, or affix markers to the target area in the second image. Variations in the markers have been described in the foregoing embodiments and will not be repeated here. The first image signal is a collection of one or more second images. For example, a second image of consecutive frames.
[0085] Please refer to Figure 11 In addition to the target area obtained by combining sound signals x1, the combined sound signals x2, ..., x... N It is also possible to obtain the same or different target regions. The processor 35 can generate markers in the second image based on these target regions (step S116) to output the first image signal y. Figure 12 For example, identifiers "1" and "3" are both target areas. Assuming they are target areas indicated by different terminal devices 10, the processor 35 adds a star-shaped marker M1 and a heart-shaped marker M2 respectively.
[0086] Processor 35 transmits the first image signal and voice signal via network 50 through communication transceiver 33 (step S105). Similarly, when terminal device 10 receives the first image signal from server 30, processor 15 can display the second image from the first image signal on display 11. At this time, one or more areas of the second image are marked. Figure 12 As shown, two regions A are labeled M1 and M2. Furthermore, the speech signal may be as follows: Figure 11 The synthesized speech signal x of multiple terminal devices 10 is shown.
[0087] In addition to indicating the selection or target area that needs attention / emphasis / marking, it is also possible to further deselect / focus / emphasize / mark. In one embodiment, the processor 35 can demark a target area in the second image to generate a second image signal. That is, unlike the first image signal, the second image signal is a second image without markings. The processor 35 can demark by removing the markings or pasting the original image of the area. Since the selection instruction generated by the terminal device 10 also includes a demark instruction, the selection instruction also corresponds to a specific sample sound signal (as a target sound signal). This target sound signal not only indicates the target area but also further indicates the demarking of the target area. Then, the processor 35 can transmit the second image signal through the communication transceiver 33 to remove the markings of the specific terminal device 10 from the second image. For example, Figure 13 This is a schematic diagram illustrating the cancellation of a mark according to an embodiment of the present invention. Please refer to... Figure 12 and Figure 13 Compared to Figure 12 Marker M1 is canceled, therefore Figure 13 M1 is not marked.
[0088] In summary, in the image marking method, terminal device, and server combining audio signals according to embodiments of the present invention, the terminal device can indicate that a target area in an image needs to be marked by combining audio signals, and the server can generate marks in the image based on the combined audio signals. Therefore, all participants can mark the shared screen, thereby improving the convenience of video conferencing and enhancing the experience of multi-person meetings.
[0089] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for image tagging incorporating sound signals, comprising: Display the first image; A selection instruction is detected, wherein the selection instruction corresponds to a target region in the first image, and the selection instruction is generated by selecting the target region through an input operation; Embedding a target sound signal into an audio signal to generate a combined sound signal, wherein the target sound signal corresponds to the target region of the selection instruction, the audio signal is obtained through audio recording, the first image is a user interface of video conferencing software, the shared screen in the user interface is divided into multiple regions, the regions including the target region, the regions respectively corresponding to multiple identifiers, the identifiers respectively corresponding to multiple sample sound signals, and the step of embedding the target sound signal into the audio signal includes: The target sound signal that matches the identifier of the target region is determined from the sample sound signals; Transmit the combined sound signal; Receive the combined sound signal; The combined sound signal is distinguished into the speech signal and the target sound signal; Determine the target region corresponding to the target sound signal; A marker is generated in the target region of the second image to generate a first image signal, wherein the first image signal includes the second image having the marker; and The first image signal and the voice signal are transmitted.
2. The image tagging method combining sound signals according to claim 1, characterized in that, The step of embedding the sample sound signal into the speech signal includes: The original sound signal is passed through a filter to generate the speech signal, wherein the original sound signal is generated by the radio, the filter is used to filter out sound signals outside the first frequency band, the speech signal belongs to the first frequency band, and the target sound signal belongs to a second frequency band higher than the first frequency band.
3. The image tagging method combining sound signals according to claim 1, characterized in that, The steps for detecting the selection instruction include: The input operation is compared with the trigger operation to generate a first comparison result; In response to the first comparison result indicating that the input operation matches the trigger operation, the target region selected by the input operation is determined; and The selection instruction is generated based on the target region.
4. The image tagging method combining sound signals according to claim 1, characterized in that, The image tagging method combining sound signals further includes: Receive the first image signal; and The second image is displayed in the first image signal, wherein the target region in the second image has the marker.
5. The image tagging method combining sound signals according to claim 4, characterized in that, The steps for detecting the selection instruction include: The input operation and the cancellation operation are compared to generate a second comparison result; In response to the second comparison result indicating that the input operation conforms to the cancellation operation, the target region selected by the input operation is determined; and The selection instruction is generated based on the target region, wherein the selection instruction is further accompanied by a cancellation instruction, and the cancellation instruction is used to remove the mark of the target region in the second image.
6. The image tagging method combining sound signals according to claim 1, characterized in that, The step of distinguishing the combined sound signal into the speech signal and the target sound signal includes: The combined audio signal is passed through a first filter to generate the speech signal, wherein the first filter is used to filter out audio signals outside the first frequency band, and the speech signal belongs to the first frequency band; and The combined sound signal is passed through a second filter to generate the target sound signal, wherein the second filter is used to filter out sound signals outside the second frequency band, the target sound signal belongs to the second frequency band, and the second frequency band is higher than the first frequency band.
7. The image tagging method combining sound signals according to claim 1, characterized in that, The second image is the user interface of the video conferencing software. The shared screen in the user interface is divided into multiple regions, including the target region. Each region corresponds to a multiple identifier, and each identifier corresponds to a multiple sample audio signal. The step of determining the target region corresponding to the target audio signal includes: Determine that the first sample sound signal in the sample sound signal matches the target sound signal; and The target region is determined based on the identifier of the first sample sound signal.
8. The image tagging method combining sound signals according to claim 1, characterized in that, The image tagging method combining sound signals further includes: The marker is removed from the target region in the second image to generate a second image signal, wherein the second image signal does not contain the marker in the second image; and The second image signal is transmitted.
9. A terminal device, characterized in that, The terminal device includes: monitor; Communication transceiver; Input devices; Memory, used to store program code; and The processor, coupled to the display, the transceiver, the input device, and the memory, is configured to load the program code for execution. The first image is displayed on the monitor; A selection instruction is detected, wherein the selection instruction corresponds to a target region in the first image, and the selection instruction is generated by selecting the target region through an input operation received by the input device; The target sound signal is embedded into the speech signal to generate a combined sound signal, wherein the target sound signal corresponds to the target region of the selection instruction, the speech signal is obtained through audio recording, the first image is a user interface of video conferencing software, the shared screen in the user interface is divided into multiple regions, the regions including the target region, the regions respectively corresponding to multiple identifiers, the identifiers respectively corresponding to multiple sample sound signals, and the processor is further configured to: The target sound signal that matches the identifier of the target region is determined from the sample sound signals; The combined audio signal is transmitted via the communication transceiver; A marker is generated in the target region of the second image to generate a first image signal, wherein the first image signal includes the second image having the marker; and The first image signal and the voice signal are transmitted through the communication transceiver.
10. The terminal device according to claim 9, characterized in that, The processor is also configured to: The original sound signal is passed through a filter to generate the speech signal, wherein the original sound signal is generated by the radio, the filter is used to filter out sound signals outside the first frequency band, the speech signal belongs to the first frequency band, and the target sound signal belongs to a second frequency band higher than the first frequency band.
11. The terminal device according to claim 9, characterized in that, The processor is also configured to: The input operation is compared with the trigger operation to generate a first comparison result; If the first comparison result indicates that the input operation matches the triggering operation, the target region selected by the input operation is determined. as well as The selection instruction is generated based on the target region.
12. The terminal device according to claim 9, characterized in that, The processor is also configured to: Receive the first image signal through the communication transceiver; and The second image in the first image signal is displayed on the display, wherein the target region in the second image has the mark.
13. The terminal device according to claim 12, characterized in that, The processor is also configured to: The input operation and the cancellation operation are compared to generate a second comparison result; If the second comparison result indicates that the input operation conforms to the cancellation operation, the target region selected by the input operation is determined. as well as The selection instruction is generated based on the target region, wherein the selection instruction is further accompanied by a cancellation instruction, and the cancellation instruction is used to remove the mark of the target region in the second image.
14. A server, characterized in that, The server includes: Communication transceiver; Memory, used to store program code; and The processor, coupled to the transceiver and the memory, is configured to load the program code for execution. Receive combined audio signals via the communication transceiver; The combined sound signal is distinguished into a speech signal and a target sound signal, wherein the speech signal is obtained by sound recording; Determine the target region corresponding to the target sound signal; A marker is generated in the target region of the second image to generate a first image signal, wherein the first image signal includes the second image having the marker, the second image being a user interface of video conferencing software, the shared screen in the user interface being divided into multiple regions, the regions including the target region, each region corresponding to a multiple identifier, each identifier corresponding to a multiple sample audio signal, and the processor is further configured to: Determine that the first sample sound signal in the sample sound signal matches the target sound signal; and The target region is determined based on the identifier of the first sample sound signal; and The first image signal and the voice signal are transmitted through the communication transceiver.
15. The server according to claim 14, characterized in that, The processor is also configured to: The combined sound signal is passed through a first filter to generate the speech signal, wherein the first filter is used to filter out sound signals outside the first frequency band, and the speech signal belongs to the first frequency band; as well as The combined sound signal is passed through a second filter to generate the target sound signal, wherein the second filter is used to filter out sound signals outside the second frequency band, the target sound signal belongs to the second frequency band, and the second frequency band is higher than the first frequency band.
16. The server according to claim 14, characterized in that, The processor is also configured to: The mark is removed from the target region in the second image to generate a second image signal, wherein the second image signal does not contain the mark of the second image; as well as The second image signal is transmitted via the communication transceiver.
Citation Information
Patent Citations
CN106575361A
EP2579588A2