Electronic device and method for adjusting display area based on action

The processor automatically detects the image nodes and generates adjustment instructions, and dynamically adjusts the display area, solving the problem that users cannot clearly see the coach's movements while exercising, improving the fitness experience.

CN120568142APending Publication Date: 2025-08-29WISTRON CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410305018.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-02-27
Filing Date
2024-03-18
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

When users use their mobile phone to watch fitness demonstration videos, they cannot clearly see the coach's movement details, and frequent operation of the mobile phone affects the exercise experience.

Method used

The processor automatically detects the target object node in the image, calculates scores and generates adjustment instructions, dynamically adjusts the display area, including enlargement, reduction and translation, and judges important display areas based on audio and image characteristics.

Benefits of technology

The image details can be clearly displayed without manual operation by the user, improving the sports experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120568142A_ABST
    Figure CN120568142A_ABST
Patent Text Reader

Abstract

The invention discloses an electronic device and method for adjusting a display area based on an action. The method comprises the following steps: acquiring a first image containing a first target object; detecting the first image to obtain a first node of the first target object; calculating a first score according to the first displacement of the first node; generating an adjustment instruction for adjusting a first display area of the first image according to the first node in response to the fact that the first score is greater than a threshold value; and outputting an adjustment instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing technology, and in particular to an electronic device and method for adjusting a display area based on motion. Background Art

[0002] Currently, many fitness apps offer demonstration videos. Users can play these videos on their phones and follow the instructor's movements. However, due to the small screen size of mobile phones, users often cannot clearly see the instructor's movements. While users can zoom in on specific parts of the image by operating their phones, frequent manipulation can affect their exercise experience. Summary of the Invention

[0003] The invention provides an electronic device and method for adjusting a display area based on an action, which can automatically adjust the display area of ​​a played image.

[0004] An electronic device for adjusting a display area based on motion according to the present invention includes a processor and a transceiver. The transceiver obtains a first image including a first target object. The processor is coupled to the transceiver and configured to: detect the first image to obtain a first node of the first target object; calculate a first score based on a first displacement of the first node; in response to the first score being greater than a threshold, generate an adjustment instruction for adjusting a first display area of ​​the first image based on the first node; and output the adjustment instruction via the transceiver.

[0005] In one embodiment of the present invention, the processor is further configured to: calculate a second score based on a second displacement of a second node of the first object; and generate an adjustment instruction based on the first node and the second node in response to the second score being greater than a threshold.

[0006] In one embodiment of the present invention, the above-mentioned processor is further configured to perform: segmenting the area of ​​interest of the first image to obtain multiple sub-regions, wherein the multiple sub-regions include a first sub-region and a second sub-region; determining based on the first image that the first node is located in the first sub-region and the second node is located in the second sub-region; and in response to the second sub-region being adjacent to the first sub-region, generating an adjustment instruction based on the first node and the second node.

[0007] In one embodiment of the present invention, the above-mentioned multiple sub-regions include a third sub-region corresponding to a third node of the first target object, wherein a third score corresponding to the third node is greater than a threshold, and the processor is further configured to execute: in response to the first sub-region being adjacent to the second sub-region and the third sub-region being adjacent to at least one of the first sub-region and the second sub-region, generate an adjustment instruction according to the first node, the second node and the third node.

[0008] In one embodiment of the present invention, the processor is further configured to execute: generating a first display area based on at least one sub-area including the first sub-area; and generating a second display area of ​​the first image based on the second sub-area in response to the second sub-area being not adjacent to the at least one sub-area.

[0009] In one embodiment of the present invention, the adjustment instruction instructs the first display area of ​​the first image to be output during a first period of time and the second display area of ​​the first image to be output during a second period of time.

[0010] In one embodiment of the present invention, the processor is further configured to generate an adjustment instruction including a zoom-in operation or a zoom-out operation in response to a size of the second display area being different from a size of the first display area.

[0011] In one embodiment of the present invention, the processor is further configured to execute: determining a boundary of the first display area according to a union of the first sub-area and the second sub-area.

[0012] In one embodiment of the present invention, the processor is further configured to: obtain a first audio file corresponding to the first image through the transceiver; determine a first correlation between the first audio file and the first node; and calculate a first score based on the first displacement and the first correlation.

[0013] In one embodiment of the present invention, the processor is further configured to perform: performing speech-to-text conversion on the first audio file to generate text; determining the number of words associated with the first node in the text; and determining the first relevance based on the number of words.

[0014] In one embodiment of the present invention, the processor is further configured to: determine a first weight of the first displacement and a second weight of the first correlation according to the category of the first image; and calculate a first score according to the first displacement, the first weight, the first correlation and the second weight.

[0015] In one embodiment of the present invention, the processor is further configured to: obtain a second image including a second object according to the transceiver; and generate a playback control instruction of the first image according to the second image, and output the playback control instruction through the transceiver.

[0016] In one embodiment of the present invention, the processor is further configured to perform: detecting a second image to obtain a plurality of limb nodes of a second target object, wherein the plurality of limb nodes include a left elbow node, a right elbow node, a left knee node, and a right knee node; calculating an angle based on the plurality of limb nodes, and determining whether the angle is within a preset range; and generating a playback control instruction for pausing the first image in response to the angle exceeding the preset range.

[0017] In one embodiment of the present invention, the processor is further configured to perform: performing facial recognition on a second target object in a second image to obtain a plurality of nodes; calculating an area change of a polygon formed by the plurality of nodes; and generating a playback control instruction for playback speed control in response to an absolute value of the area change being greater than a change threshold.

[0018] In one embodiment of the present invention, the playback control instruction is used to reduce the playback speed of the first image.

[0019] In one embodiment of the present invention, the processor is further configured to perform: performing facial recognition on the second target in the second image to obtain a plurality of nodes; calculating an area change of a polygon formed by the plurality of nodes; and generating a playback control instruction for playback speed control based on the area change.

[0020] In one embodiment of the present invention, the above-mentioned processor is further configured to perform: in response to the area change being greater than a first threshold, generating a playback control instruction for reducing the playback speed of the first image; and in response to the area change being less than or equal to a second threshold, generating a playback control instruction for increasing the playback speed of the first image.

[0021] In one embodiment of the present invention, the processor is further configured to: detect the first image to obtain the center point and height of the first target; and determine the region of interest of the first image according to the center point and height.

[0022] In one embodiment of the present invention, the processor is further configured to execute: outputting a script via a transceiver, wherein the script includes an adjustment instruction and a timestamp of the first image corresponding to the adjustment instruction.

[0023] A method for adjusting a display area based on motion according to the present invention comprises: obtaining a first image including a first target object; detecting the first image to obtain a first node of the first target object; calculating a first score based on a first displacement of the first node; generating an adjustment instruction for adjusting a first display area of ​​the first image based on the first node in response to the first score being greater than a threshold; and outputting the adjustment instruction.

[0024] Based on the above, the electronic device of the present invention can adjust the display area of ​​the image according to the movements of the characters in the image, so that the local details of the image can be presented to the user more clearly. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 A schematic diagram of an electronic device that adjusts a display area based on motion is shown as an embodiment of the present invention;

[0026] Figure 2 A schematic diagram illustrating the generation and application of a script according to an embodiment of the present invention is provided;

[0027] Figure 3 A schematic diagram illustrating a region of interest of an exemplary image according to an embodiment of the present invention;

[0028] Figure 4 A schematic diagram illustrating adjustment of a display area according to an embodiment of the present invention;

[0029] Figure 5 A schematic diagram illustrating adjustment of a display area according to an embodiment of the present invention;

[0030] Figure 6 A schematic diagram illustrating a target object according to an embodiment of the present invention;

[0031] Figure 7 A flow chart of a method for adjusting a display area based on motion is shown according to an embodiment of the present invention.

[0032] Explanation of symbols

[0033] 100: Electronic devices

[0034] 110: Processor

[0035] 120: Storage Media

[0036] 130: transceiver

[0037] 20,30: Target

[0038] 21,22,23,24,31,32,33,34,41,42,43,44: nodes

[0039] 300,400:Terminal device

[0040] 35: Face

[0041] 40: Polygon

[0042] 50: Region of interest

[0043] 500: Cloud Server

[0044] 51,52,53,54: sub-areas

[0045] 510,520,530,540: display area

[0046] A, B: weight

[0047] C: Center point

[0048] H: Height

[0049] S701, S702, S703, S704, S705: Steps DETAILED DESCRIPTION

[0050] When a user is exercising while watching a coach's demonstration video, if they cannot clearly see the details of the demonstration video, they must manually adjust the display area of ​​the demonstration video using their terminal device. This affects the user's exercise experience. To address this issue, the present invention provides a method for automatically adjusting the display area of ​​an image.

[0051] Figure 1 A schematic diagram of an electronic device 100 for adjusting a display area based on motion is shown according to an embodiment of the present invention. The electronic device 100 may include a processor 110 , a storage medium 120 , and a transceiver 130 .

[0052] The processor 110 may be, for example, a central processing unit (CPU), or other programmable general-purpose or special-purpose microcontrol unit (MCU), microprocessor, digital signal processor (DSP), programmable controller, application-specific integrated circuit (ASIC), graphics processing unit (GPU), image signal processor (ISP), image processing unit (IPU), arithmetic logic unit (ALU), complex programmable logic device (CPLD), field programmable gate array (FPGA), or other similar components or combinations thereof. The processor 110 may be coupled to the storage medium 120 and the transceiver 130 to access and execute multiple modules and various applications stored in the storage medium 120.

[0053] The storage medium 120 is, for example, any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk drive (HDD), solid state drive (SSD), or similar elements or a combination of the above elements, and is used to store multiple modules or various applications that can be executed by the processor 110.

[0054] The transceiver 130 transmits or receives signals wirelessly or by wire. The transceiver 130 may also perform operations such as low-noise amplification, impedance matching, frequency mixing, up- or down-conversion of frequencies, filtering, amplification, and the like. The processor 110 may be communicatively coupled to an external electronic device via the transceiver 130.

[0055] Figure 2 A schematic diagram illustrating the generation and application of a script according to an embodiment of the present invention is shown. The electronic device 100 can be communicatively connected to the cloud server 500 and one or more terminal devices (eg, the terminal device 300 or the terminal device 400 ) via the transceiver 130 .

[0056] The user can operate the terminal device 300 to play a video containing the target object 20 (such as Figure 3 , where the target 20 is, for example, a trainer teaching a user how to exercise. Furthermore, the terminal device 300 can capture a user image containing the target 30, where the target 30 is, for example, a user exercising (i.e., the owner of the terminal device 300). The terminal device 300 can transmit the user image to the electronic device 100.

[0057] The electronic device 100 can receive a demonstration image (e.g., a demonstration image downloaded from the internet, such as an aerobic dance demonstration video or a yoga instructional video) containing a target object 20 (e.g., a trainer) or a user image containing a target object 30 (e.g., a user) via the transceiver 130. The electronic device 100 can generate adjustment instructions based on the demonstration image or generate playback control instructions based on the user image. The electronic device 100 can transmit the adjustment instructions or playback control instructions to the terminal device 300 via the transceiver 130. The terminal device 300 can adjust the display area of ​​the demonstration image based on the adjustment instructions or control the playback of the demonstration image based on the playback control instructions.

[0058] In one embodiment, the adjustment command may include but is not limited to a zoom-in command, a zoom-out command, or a pan command.

[0059] In one embodiment, the playback control command may include but is not limited to a fast forward command, a rewind command, a slow motion command, a pause command, or a (resume) play command.

[0060] Figure 3 According to one embodiment of the present invention, a schematic diagram of a region of interest 50 of a demonstration image is shown. After receiving a demonstration image including a target object 20 (e.g., a coach), the processor 110 may perform object detection to detect the demonstration image, thereby obtaining the center point C and height H of the target object 20 (e.g., the coach). The processor 110 may determine the region of interest 50 of the demonstration image based on the center point C and height H of the target object 20 (e.g., the coach). For example, the shortest distance between the upper boundary (or lower boundary) of the region of interest 50 and the target object 20 (e.g., the coach) may be B×H, where B may be a preset weight. The distance between the left boundary (or right boundary) of the region of interest 50 and the center point C may be A×H, where A may be a preset weight.

[0061] The processor 110 may perform object detection to detect the region of interest 50 of the demonstration image, thereby obtaining one or more joints of the target object 20 (e.g., a coach). For example, the processor 110 may obtain a node 21 corresponding to the right palm of the target object 20 (e.g., a coach), a node 22 corresponding to the right elbow of the target object 20 (e.g., a coach), a node 23 corresponding to the left palm of the target object 20 (e.g., a coach), or a node 24 corresponding to the left sole of the target object 20 (e.g., a coach). It should be noted that the target object 20 (e.g., a coach) may have a plurality of joints not in the target object 20 (e.g., a coach). Figures 3 to 5 One or more nodes are shown.

[0062] The processor 110 may calculate a score corresponding to the node based on the displacement of the node. In one embodiment, the displacement of the node may be proportional to the score of the node. That is, a node with a larger displacement may have a higher score.

[0063] In one embodiment, the score of a node may be associated with an audio file synchronized with the demonstration image. The processor 110 may obtain the audio file synchronized with the demonstration image through the transceiver 130. The processor 110 may determine the correlation between the audio file and the node, and may calculate the score of the node based on the displacement of the node and the correlation of the node. Specifically, the processor 110 may perform speech to text (STT) on the audio file to generate text. Then, the processor 110 may determine the number of words associated with the node in the text, and then determine the correlation between the node and the audio file based on the number of associated words. For example, if the coach in the demonstration video continuously repeats the instruction of "rotate your right palm", the processor 110 may determine that the audio file has a high correlation with node 21 based on the multiple occurrences of the word "right palm" in the text of the audio file.

[0064] In one embodiment, the storage medium 120 may pre-store a lookup table, wherein the lookup table may include a mapping relationship between one or more words and a preset word. Assume that a plurality of words having a mapping relationship with the preset word appear in the text of the audio file. The processor 110 may determine the relevance between the node and the audio file based on the number of the plurality of words. For example, the lookup table may record the mapping relationship between words such as "buttocks", "coccyx" or "sit bone" and the preset word "buttocks". If the text in the audio file contains "coccyx" once and "sit bone" twice, the processor 110 may determine the relevance between the audio file and the node corresponding to "buttocks" based on the number "3".

[0065] When calculating the score of a node based on the displacement and relevance of the node, the displacement and the relevance may have respective weights. The processor 110 may calculate the score of the node based on the displacement, the weight of the displacement, the relevance, and the weight of the relevance. In one embodiment, the processor 110 may determine the weight of the displacement or the weight of the relevance based on the category of the demonstration image. For example, if the theme of the demonstration image is fitness guidance, the coach in the representative demonstration image may often give instructions to the user by voice. Accordingly, the processor 110 may increase the weight of the relevance of the node and reduce the weight of the displacement of the node to enhance the influence of the audio file on the calculation of the node score.

[0066] The processor 110 may determine, based on the score of the node, whether the node can be used to generate an adjustment instruction for the display area of ​​the demonstration image. If the score corresponding to the node is greater than a threshold, the processor 110 may determine that the node can be used to generate an adjustment instruction. If the score corresponding to the node is less than or equal to the threshold, the processor 110 may determine that the node cannot be used to generate an adjustment instruction. For example, if the score of node 21 is greater than the threshold and the score of node 22 is less than or equal to the threshold, the processor 110 may generate an adjustment instruction based on node 21 but not based on node 22. For another example, if the score of node 21 is greater than the threshold and the score of node 22 is also greater than the threshold, the processor 110 may generate an adjustment instruction based on node 21 and node 22.

[0067] Processor 110 may segment region of interest 50 of the exemplary image to obtain multiple subregions (e.g., a grid) and determine the subregion in which each node is currently located. For example, processor 110 may determine that node 21 is located in subregion 51, node 22 is located in subregion 52, node 23 is located in subregion 53, or node 24 is located in subregion 54. Subregions containing nodes with scores greater than a threshold are hereinafter referred to as selected subregions.

[0068] Figure 4 A schematic diagram illustrating adjustment of a display area is shown according to one embodiment of the present invention. The processor 110 may generate a corresponding display area based on adjacent selected sub-areas. If multiple selected sub-areas are used to generate adjustment instructions for a display area, one of the multiple selected sub-areas must be adjacent to one or more other sub-areas in the multiple selected sub-areas. For example, assume that sub-areas 51, 52, 53, and 54 are all selected sub-areas. The processor 110 may determine that sub-areas 51 and 52 can be used to generate adjustment instructions for display area 510 based on the fact that sub-area 52 is adjacent to sub-area 51, where display area 510 corresponds to timestamp (N-1). Furthermore, the processor 110 may determine that sub-area 53 can also be used to generate adjustment instructions for display area 510 based on the fact that sub-area 53 is adjacent to sub-area 51. Since sub-area 54 is not adjacent to any of sub-areas 51, 52, or 53, the processor 110 may determine that sub-area 54 cannot be used to generate adjustment instructions for display area 510.

[0069] In one embodiment, if multiple selected sub-regions correspond to the display area, the processor 110 may determine the boundary of the display area based on the union of the multiple selected sub-regions. For example, the processor 110 may determine the boundary of the display area 510 based on the union of the multiple selected sub-regions including sub-regions 51, 52, and 53. Since sub-region 51 or sub-region 52 is located at the upper boundary of the union, the processor 110 may determine the upper boundary of the display area 510 based on sub-region 51 or sub-region 52.

[0070] When the timestamp of the demonstration image advances from (N-1) to (N), processor 110 may generate a new display area 520 for the demonstration image corresponding to timestamp (N). The multiple selected sub-areas corresponding to display area 520 include, for example, sub-areas 51, 52, 53, and 54. Processor 110 may generate adjustment instructions based on display area 510 corresponding to timestamp (N-1) and display area 520 corresponding to timestamp (N). For example, processor 110 may generate a zoom-out instruction based on the fact that the size of display area 520 is larger than the size of display area 510, so that the display of terminal device 300 can fully display the entire display area 520. Furthermore, processor 110 may generate a translation instruction based on the fact that the center point of display area 520 is different from the center point of display area 510.

[0071] The processor 110 may generate multiple display areas of the exemplary image within the same time period, and a selected sub-area of ​​one display area is not adjacent to a selected sub-area of ​​another display area. Figure 5 A schematic diagram illustrating adjustment of a display area according to an embodiment of the present invention is shown. Processor 110 may generate display area 530 based on a plurality of selected sub-areas including sub-areas 51, 52, and 53, and may generate display area 540 based on a plurality of selected sub-areas including sub-area 54, wherein a selected sub-area of ​​display area 540 (e.g., sub-area 54) is not adjacent to a selected sub-area of ​​display area 540 (e.g., sub-area 51, 52, or 53).

[0072] Processor 110 may generate adjustment instructions based on the multiple display regions of the exemplary image. The adjustment instructions may be used to instruct a terminal device (e.g., terminal device 300) to output different display regions during different time periods. For example, the adjustment instructions may instruct terminal device 300 to output the image in display region 530 during time period (M-1) and to output the image in display region 540 during time period (M).

[0073] In one embodiment, processor 110 may generate a translation instruction based on the fact that the multiple display areas of the exemplary image have different center points. For example, processor 110 may generate a translation instruction based on the fact that the center point of display area 540 is different from the center point of display area 530. Terminal device 300 may translate the output image from display area 530 to display area 540 according to the translation instruction.

[0074] In one embodiment, processor 110 may generate an adjustment instruction for a zoom-in or zoom-out operation based on the different sizes of the multiple display areas of the demonstration image. For example, processor 110 may generate a zoom-out instruction based on the fact that display area 540 is larger than display area 530. Terminal device 300 may perform a zoom-out operation on the output image in accordance with the zoom-out instruction. Consequently, the user may view a smaller portion of the demonstration image (with clearer details or higher resolution) during time period (M-1) and a larger portion of the demonstration image (with less clear details or lower resolution) during time period (M).

[0075] The processor 110 may generate a playback control instruction according to a user image including the target object 30 (ie, the user of the terminal device 300 ). Figure 6 A schematic diagram of a target object 30 (e.g., a user) is shown according to one embodiment of the present invention. In one embodiment, the processor 110 can determine, based on the user image, whether the target object 30 (e.g., the user) is still looking at the terminal device 300 or has moved away from the terminal device 300, and thus determine whether to instruct the terminal device 300 to pause playback of the demonstration image. Specifically, the processor 110 can perform object detection on the user image to obtain multiple limb nodes of the target object 30 (e.g., the user). The multiple limb nodes can include a node 31 representing the left elbow, a node 32 representing the right elbow, a node 33 representing the left knee, and a node 34 representing the right knee. The processor 110 can calculate angles based on the multiple limb nodes and determine whether the angles are within a preset range. If the angles are within the preset range, the processor 110 can determine that the target object 30 (e.g., the user) is still in front of the terminal device 300 (i.e., the target object 30 is still visible in the user image captured by the terminal device 300). Accordingly, the processor 110 may not instruct the terminal device 300 to pause playback of the demonstration image. If the angle exceeds the preset range, the processor 110 may determine that the target object 30 has moved away from the terminal device 300 (i.e., the target object 30 does not appear in the user image captured by the terminal device 300). Based on this, the processor 110 may generate a pause instruction, wherein the pause instruction may instruct the terminal device 300 to pause playing the demonstration image. The processor 110 may calculate the angle θ according to formula (1) and formula (2): hand and angle θ leg , and the angle θ can be determined according to formula (3) hand or angle θ leg Whether it exceeds the preset range, where (x1, y1) is the coordinate of node 31, (x2, y2) is the coordinate of node 32, (x3, y3) is the coordinate of node 33, and (x4, y4) is the coordinate of node 34. If formula (3) does not hold, processor 110 may generate a pause instruction. The preset range of 10 to 180 degrees in formula (3) is determined based on multiple image samples of limb nodes of different people, for example.

[0076]

[0077]

[0078] 10°≤θ hand ,θ leg ≤180°…(3)

[0079] The processor 110 can determine whether to generate a playback control instruction for playback speed control (e.g., fast forward, rewind, or slow motion) based on the expression of the target object 30 (e.g., the user). Specifically, the processor 110 can perform facial recognition on the face 35 of the target object 30 (e.g., the user) in the user image to obtain multiple nodes, such as a node 41 representing the left eye, a node 42 representing the right eye, a node 43 representing the left corner of the mouth, and a node 44 representing the right corner of the mouth. The processor 110 can calculate the area change of the polygon 40 formed by the multiple nodes. If the absolute value of the area change is greater than a change threshold for a time greater than a time threshold, it indicates that the target object 30 (e.g., the user) may be surprised or confused by the demonstration image. Based on this, the processor 110 can generate a playback control instruction for reducing the playback speed (e.g., a slow motion instruction). If the absolute value of the area change is less than or equal to the change threshold, or the absolute value of the area change is greater than the change threshold for a time less than or equal to a time threshold, the processor 110 may not generate a playback control instruction for playback speed control. The processor 110 can calculate the area P of the polygon 40 at the time point t according to formula (4): t , and the area change ΔP of polygon 40 during the time period (t2-t1) can be calculated according to formula (5), where (x1, y1) represents the coordinates of node 31, (x2, y2) represents the coordinates of node 32, (x3, y3) represents the coordinates of node 33, and (x4, y4) represents the coordinates of node 34.

[0080]

[0081] ΔP=P t2 -P t1 …(5)

[0082] In one embodiment, the processor 110 can determine whether the face of the target object 30 (e.g., the user) is approaching or moving away from the playback device (e.g., the terminal device 300) playing the demonstration video based on the area change ΔP of the polygon 40. If the area change ΔP is greater than a threshold, it means that the target object 30 (e.g., the user) may not be able to see the demonstration video clearly, and the target object 30 (e.g., the user) needs to move their face closer to the playback device. Based on this, the processor 110 can generate a playback control instruction (e.g., a slow motion instruction) to reduce the playback speed, making it easier for the user to see the details of the demonstration video. If the area change ΔP is less than or equal to the threshold, it means that the target object 30 (e.g., the user) does not have difficulty seeing the demonstration video and therefore has not moved their face closer to the playback device. Based on this, the processor 110 can generate a playback control instruction to increase (or restore) the playback speed.

[0083] In one embodiment, the processor 110 may perform facial expression recognition on the face 35 of the target object 30 (e.g., the user) in the user image based on, for example, machine learning techniques, and generate a playback control instruction for controlling the playback speed (e.g., fast forward, rewind, or slow motion) based on the recognition result of the facial expression recognition. For example, if the recognition result of the facial expression recognition indicates that the user is confused, the processor 110 may generate a playback control instruction for reducing the playback speed (e.g., a slow motion instruction) to make it easier for the user to see the details of the demonstration video.

[0084] Reference Figure 2 In one embodiment, after generating an adjustment instruction or a playback control instruction, the processor 110 may generate a script corresponding to the demonstration image, wherein the script may include the adjustment instruction or the playback control instruction and may include a timestamp corresponding to the adjustment instruction or the playback control instruction. The processor 110 may upload the script to the cloud server 500 for other users to download. For example, the terminal device 400 may download the script from the cloud server 500. The terminal device 400 may execute the script when playing the demonstration image to adjust the display area or playback mode of the demonstration image. Accordingly, the terminal device 400 can automatically adjust the playback of the demonstration image without the need for data transmission with the electronic device 100.

[0085] Figure 7 According to an embodiment of the present invention, a flow chart of a method for adjusting a display area based on motion is shown, wherein the method may be performed as follows: Figure 1 The electronic device 100 shown in FIG. 1 is implemented as follows. In step S701, a first image including a first target object is obtained. In step S702, the first image is detected to obtain a first node of the first target object. In step S703, a first score is calculated based on a first displacement of the first node. In step S704, in response to the first score being greater than a threshold, an adjustment instruction is generated based on the first node for adjusting a first display area of ​​the first image. In step S705, the adjustment instruction is output.

[0086] In summary, the electronic device of the present invention can determine which area of ​​the image is the important display area of ​​user interest based on the movements of the characters in the image being played. The electronic device can further determine the important display area based on the audio of the image. After determining the important display area, the electronic device can output instructions based on the important display area to adjust the playback mode of the image played by the user's terminal device. The present invention provides users with a convenient image playback method. Users do not need to manually operate the terminal device to change the display area of ​​the image being played.

Claims

1. An electronic device for adjusting a display area based on an action, comprising: a transceiver for acquiring a first image including a first target object; as well as a processor coupled to the transceiver and configured to execute: detecting the first image to obtain a first node of the first target; calculating a first score based on a first displacement of the first node; In response to the first score being greater than a threshold, generating, according to the first node, an adjustment instruction for adjusting a first display area of ​​the first image; and The adjustment instruction is outputted through the transceiver.

2. The electronic device of claim 1 , wherein the processor is further configured to perform: calculating a second score based on a second displacement of a second node of the first object; and In response to the second score being greater than the threshold, the adjustment instruction is generated based on the first node and the second node.

3. The electronic device of claim 2 , wherein the processor is further configured to perform: Segmenting the region of interest of the first image to obtain a plurality of sub-regions, wherein the plurality of sub-regions include a first sub-region and a second sub-region; determining, based on the first image, that the first node is located in the first sub-region and the second node is located in the second sub-region; and In response to the second sub-region being adjacent to the first sub-region, the adjustment instruction is generated according to the first node and the second node.

4. The electronic device of claim 3 , wherein the plurality of sub-regions include a third sub-region corresponding to a third node of the first object, wherein a third score corresponding to the third node is greater than the threshold, and the processor is further configured to execute: In response to the first sub-region being adjacent to the second sub-region and the third sub-region being adjacent to at least one of the first sub-region and the second sub-region, the adjustment instruction is generated according to the first node, the second node, and the third node.

5. The electronic device of claim 3 , wherein the processor is further configured to perform: generating the first display area according to at least one sub-area including the first sub-area; and In response to the second sub-region being not adjacent to the at least one sub-region, a second display region of the first image is generated according to the second sub-region. 6 . The electronic device of claim 5 , wherein the adjustment instruction instructs the first display area of ​​the first image to be output during a first time period and the second display area of ​​the first image to be output during a second time period.

7. The electronic device of claim 6, wherein the processor is further configured to perform: In response to the size of the second display area being different from the size of the first display area, the adjustment instruction including a zoom-in operation or a zoom-out operation is generated.

8. The electronic device of claim 3 , wherein the processor is further configured to perform: The boundary of the first display area is determined according to a union of the first sub-area and the second sub-area.

9. The electronic device of claim 1 , wherein the processor is further configured to perform: obtaining, through the transceiver, a first audio file corresponding to the first image; Determining a first correlation between the first audio file and the first node; and The first score is calculated based on the first displacement and the first correlation.

10. The electronic device of claim 9, wherein the processor is further configured to perform: performing speech-to-text conversion on the first audio file to generate text; determining the number of words in the text associated with the first node; and The first relevance is determined based on the number of the words.

11. The electronic device of claim 9, wherein the processor is further configured to perform: determining a first weight of the first displacement and a second weight of the first correlation according to the category of the first image; and The first score is calculated based on the first displacement, the first weight, the first correlation, and the second weight.

12. The electronic device of claim 1 , wherein the processor is further configured to perform: Acquiring a second image including a second target object according to the transceiver; and A playback control instruction for the first image is generated according to the second image, and the playback control instruction is output through the transceiver.

13. The electronic device of claim 12, wherein the processor is further configured to perform: detecting the second image to obtain a plurality of limb nodes of the second target, wherein the plurality of limb nodes include a left elbow node, a right elbow node, a left knee node, and a right knee node; Calculating an angle based on the plurality of limb nodes, and determining that the angle is within a preset range; as well as In response to the angle exceeding the preset range, the playback control instruction for pausing the first image is generated.

14. The electronic device of claim 1 , wherein the processor is further configured to perform: performing facial recognition on the second target in the second image to obtain a plurality of nodes; calculating a change in area of ​​a polygon formed by the plurality of nodes; and In response to the absolute value of the area change being greater than a change threshold, the playback control instruction for playback speed control is generated. 15 . The electronic device as claimed in claim 14 , wherein the playback control instruction is used to reduce the playback speed of the first image.

16. The electronic device of claim 1 , wherein the processor is further configured to perform: performing facial recognition on the second target in the second image to obtain a plurality of nodes; calculating a change in area of ​​a polygon formed by the plurality of nodes; and The playback control instruction for playback speed control is generated according to the area change.

17. The electronic device of claim 16, wherein the processor is further configured to perform: In response to the area change being greater than a first threshold, generating the playback control instruction for reducing the playback speed of the first image; and In response to the area change being less than or equal to a second threshold, generating the playback control instruction for increasing the playback speed of the first image.

18. The electronic device of claim 1 , wherein the processor is further configured to perform: detecting the first image to obtain the center point and height of the first target; and The region of interest of the first image is determined according to the center point and the height.

19. The electronic device of claim 1 , wherein the processor is further configured to perform: A script is outputted through the transceiver, wherein the script includes the adjustment instruction and a timestamp of the first image corresponding to the adjustment instruction.

20. A method for adjusting a display area based on an action, comprising: Acquire a first image including a first target object; detecting the first image to obtain a first node of the first target; calculating a first score based on a first displacement of the first node; In response to the first score being greater than a threshold, generating, according to the first node, an adjustment instruction for adjusting a first display area of ​​the first image; and The adjustment instruction is output.