Video generation method, device, equipment, medium and program product

By performing target detection and state machine control on candidate images in motion scenes, target videos are automatically generated, which solves the problem of inefficiency and low precision caused by manual observation and realizes efficient and high-precision automation of motion performance analysis.

CN118379479BActive Publication Date: 2025-09-16BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410294941.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-14
Publication Date
2025-09-16
Estimated Expiration
2044-03-14

AI Technical Summary

Technical Problem

Existing technologies rely on manual observation and post-data processing in sports performance analysis, resulting in low efficiency and difficulty in achieving high precision and real-time performance.

Method used

By performing target detection on candidate images, determining the positional relationship between the target area and the preset area, and generating a trigger event, the target video is automatically generated, avoiding the inefficiency and low accuracy of manual observation.

Benefits of technology

It achieves efficient automation and high precision in sports performance analysis, can accurately capture the details of athletes' movements, provide real-time feedback and detailed data support, and optimize training strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118379479B_ABST
    Figure CN118379479B_ABST
Patent Text Reader

Abstract

The present disclosure provides a video generation method, apparatus, device, medium, and program product, relating to the fields of artificial intelligence technology, specifically computer vision, deep learning, and other technical fields, and can be applied to scenarios such as video generation and smart sports. The video generation method includes: performing target detection processing on candidate images to determine a target area in the candidate images; generating a trigger event based on the positional relationship between the target area and a preset area; and in response to the trigger event, determining a target image in the candidate images and generating a target video based on the target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically computer vision, deep learning and other technical fields, and can be applied to scenarios such as video generation and smart sports. In particular, it relates to a video generation method, device, equipment, medium and program product. Background Art

[0002] With the development of artificial intelligence technology, video can be used for intelligent analysis in many fields. The demand for high-precision, real-time motion performance analysis of moving objects based on video is growing. Summary of the Invention

[0003] The present disclosure provides a video generation method, apparatus, device, and medium.

[0004] According to one aspect of the present disclosure, a video generation method is provided, comprising: performing target detection processing on a candidate image to determine a target area in the candidate image; generating a trigger event based on positional relationship information between the target area and a preset area; determining a target image in the candidate image in response to the trigger event, and generating a target video based on the target image.

[0005] According to another aspect of the present disclosure, a video generation device is provided, including: a detection module for performing target detection processing on a candidate image to determine a target area in the candidate image; a trigger module for generating a trigger event based on positional relationship information between the target area and a preset area; and a generation module for determining a target image in the candidate image in response to the trigger event, and generating a target video based on the target image.

[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute any one of the methods described in any one of the above aspects.

[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any one of the methods according to any one of the above aspects.

[0008] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of the above aspects.

[0009] The present disclosure can improve video generation efficiency and accuracy.

[0010] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0012] Figure 1 is a schematic diagram according to a first embodiment of the present disclosure;

[0013] Figure 2 is a schematic diagram of an application scenario provided according to an embodiment of the present disclosure;

[0014] Figure 3 is a schematic diagram according to a second embodiment of the present disclosure;

[0015] Figure 4 is a schematic diagram of a candidate image provided according to an embodiment of the present disclosure;

[0016] Figure 5 is a schematic diagram of the intersection of a target area and a preset area provided according to an embodiment of the present disclosure;

[0017] Figure 6 is a schematic diagram of a preset state machine provided according to an embodiment of the present disclosure;

[0018] Figure 7 is a schematic diagram of images corresponding to multiple action states provided according to an embodiment of the present disclosure;

[0019] Figure 8 is a schematic diagram according to a third embodiment of the present disclosure;

[0020] Figure 9 is a schematic diagram of an electronic device used to implement the video generation method according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0022] Related technologies often rely on manual observation and post-processing, which is not only time-consuming and labor-intensive, but also difficult to achieve the required accuracy and real-time performance. Motion performance analysis can be applied in many scenarios, such as autonomous driving, video surveillance, and sports.

[0023] With the development of computer vision technology, it is now possible to capture videos of moving objects and then analyze their performance based on them, addressing the challenges inherent in manual observation and analysis. For example, in sports, it is possible to capture target videos of athletes and conduct performance analysis based on these videos. This includes precise analysis of technical details, immediate adjustments to movement strategies, and effective management of injury prevention, thereby continuously optimizing athlete performance.

[0024] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure. This embodiment provides a video generation method, the method comprising:

[0025] 101. Perform target detection processing on a candidate image to determine a target area in the candidate image.

[0026] 102. Generate a trigger event based on the positional relationship information between the target area and the preset area.

[0027] 103. In response to the triggering event, determine a target image among the candidate images, and generate a target video based on the target image.

[0028] The candidate images are captured by an image sensor (such as a camera). For example, during the movement of a target (such as an athlete), the camera continuously captures images of the target (such as athlete images) for a preset time as candidate images.

[0029] The target region refers to the region where the target is located in the candidate image, for example, the body region of an athlete in the candidate image.

[0030] Preset areas are pre-set based on actual needs. Taking sports as an example, different preset areas can be set for different sports. There can be one or more preset areas.

[0031] After obtaining the target area and the preset area, it can be determined whether the two at least partially intersect as position relationship information.

[0032] Afterwards, a trigger event can be generated based on the positional relationship information. For example, a state machine can be used to perform a state transition based on the positional relationship information between the target area and the preset area. When the state after the transition is the start state, a trigger event is generated. Alternatively, a correspondence between positional relationship information and events can be preconfigured based on rules. If the event corresponding to the current positional relationship information in the correspondence is a trigger event, a trigger event is generated.

[0033] The target image is at least a portion of the candidate images. At least a portion of the candidate images can be selected as the target image based on a preset rule, and the target video is composed based on the target image.

[0034] Specifically, a video includes a series of images that are continuous in time. Therefore, a target video can be obtained based on the timestamp of the target image and a set frame rate.

[0035] In this embodiment, by identifying a target image from candidate images and generating a target video based on the target image, the target video can be automatically generated, thereby avoiding the inefficiency and low accuracy issues associated with manual observation. Furthermore, the target image is acquired based on a trigger event, which is generated based on the positional relationship between the target area and a preset area. This allows for accurate triggering, thereby accurately acquiring the target image and improving the accuracy of the target video.

[0036] In order to better understand the embodiments of the present disclosure, application scenarios to which the embodiments of the present disclosure can be applied are described.

[0037] Figure 2 2 is a schematic diagram of an application scenario provided according to an embodiment of the present disclosure. The scenario provides a video recording system, which includes: a camera 201, a cache memory 202, a processor 203, a controller 204, a fixed memory 205 and a player 206.

[0038] The camera 201 is used to capture candidate images of the target.

[0039] The cache memory 202 is used to cache candidate images captured by the camera. For example, the cache memory stores candidate images of a preset duration (such as m seconds), where m is a preset value.

[0040] The processor 203 is used to perform target detection processing on the candidate image to obtain the target area in the candidate image, and generate a trigger event based on the positional relationship information between the target area and the preset area. Taking a sports scene as an example, the target area can be the body area of ​​an athlete.

[0041] Controller 204 is configured to determine a target image from the candidate images in response to a trigger event and store the target image in a fixed memory. The target image can be retrieved from a specified time point after the trigger event. For example, if the trigger event occurs at time t1, a delay of m / 2 seconds can be set to time t2, and the candidate image m seconds before time t2 can be retrieved as the target image.

[0042] The fixed memory 205 is used to store the target image and compose the target video from the target image. The storage medium of the cache memory is usually volatile, such as a memory; while the storage medium of the fixed memory is non-volatile, such as a hard disk, so that the target video can be stored for a longer period of time.

[0043] The player 206 is used to play the target video. The target video can be played repeatedly as needed to better perform motion performance analysis.

[0044] In combination with the above application scenarios, the embodiments of the present disclosure further provide a video generation method.

[0045] Figure 3 is a schematic diagram according to a second embodiment of the present disclosure. This embodiment provides a video generation method, the method comprising:

[0046] 301. Obtain candidate images and cache the candidate images for a preset time period.

[0047] 302. Perform target detection processing on the candidate image to determine a target area in the candidate image.

[0048] Combine Figure 2 ,The athlete’s movement process can be captured by a camera to obtain a candidate image.

[0049] After the camera captures the candidate image, on the one hand, the candidate image can be sent to the cache memory for caching, and on the other hand, the candidate image can be sent to the processor for processing.

[0050] After the processor obtains the candidate image, it performs target detection processing on the candidate image to obtain the target area.

[0051] Specifically, the region of interest in the candidate image may be determined, and a region of interest image corresponding to the region of interest may be acquired; and target detection processing may be performed on the region of interest image to determine the target region.

[0052] In this embodiment, target detection processing is performed on the region of interest. Since the size of the image of the region of interest is usually smaller than the candidate image, this can improve processing efficiency.

[0053] Furthermore, the region of interest may be determined based on preset coordinates of a preset area.

[0054] Specifically, the coordinates of the region of interest may be calculated based on the coordinates of the preset region, thereby determining the region of interest.

[0055] See also Figure 4 , the preset areas include: area A, area B and area C, and the calculation formula of the coordinate box of the area of ​​interest E is:

[0056] box=(Center x -W,Center y -W,Center x +W, Center y +W)

[0057] W=max(C xmax -A xmin ,C ymax -A ymin )

[0058]

[0059] Among them, (Center x -W,Center y -W),(Center x +W, Center y +W) are the coordinates of the lower left corner and the upper right corner of the region of interest E;

[0060] W is the radius width of the region of interest;

[0061] Center is the coordinate of the center point of the region of interest;

[0062] C xmax is the maximum x-coordinate of the preset area C, A xmin is the minimum x-coordinate value of the preset area A, C ymax is the maximum y coordinate of the preset area C, A ymin is the minimum y-coordinate value of the preset area A; the coordinate of the preset area is the preset value.

[0063] In this embodiment, the region of interest is determined based on the preset coordinates of the preset area, so that a precise region of interest can be obtained, thereby improving processing accuracy.

[0064] After determining the region of interest image, target detection processing is performed on the region of interest image to obtain the target region.

[0065] Combine Figure 4 or Figure 5 , the target area is the athlete's body area D.

[0066] 303. Use the positional relationship information between the target area and the preset area as a current transition condition, and transition from a first state of a preset state machine to a second state according to the current transition condition; the preset state of the preset state machine includes: a startup state.

[0067] 304. If the second state is the start state, generate a trigger event.

[0068] In this embodiment, the trigger event is generated based on the state machine, which can improve the accuracy of the trigger event and the accuracy of generating the target video.

[0069] The first state is a pre-transfer state of the preset state machine, and the second state is a post-transfer state of the preset state machine.

[0070] The startup state is a preset state of the preset state machine, and the preset startup state corresponds to a trigger event, that is, when the preset state of the preset state machine is transferred to the startup state, a trigger event is generated.

[0071] Furthermore, there are multiple preset areas; the position relationship information is used to indicate that the target area has an intersection relationship with at least one preset area among the multiple preset areas; the preset state also includes: a no-action state; using the position relationship information between the target area and the preset area as the current transfer condition, and transferring from the first state of the preset state machine to the second state according to the current transfer condition, including: using the position relationship information as each current transfer condition respectively, starting from the first state being the no-action state, and successively determining the target state of the first state under each current transfer condition as the second state.

[0072] The no-action state is also a preset state of the preset state machine, and the no-action state is the starting state of the preset state machine, that is, the preset state machine starts to perform state transfer operations from the no-action state.

[0073] In this embodiment, the first state starts from the no-action state, and the second state is determined in sequence according to the current transition condition. The second state can be gradually determined until the second state is the start state and the trigger condition is generated, thereby improving processing accuracy.

[0074] In some embodiments, the target area is a body area of ​​the subject when performing a target sport; the target sport includes a plurality of preset key actions to be performed in sequence;

[0075] Each of the plurality of preset areas corresponds to each of the plurality of preset key actions;

[0076] The preset state further includes: a plurality of action states, and each action state in the plurality of action states corresponds one-to-one to each of the preset key actions.

[0077] For example, the target object is an athlete, the target sport is vaulting, and the multiple preset key actions are: take-off action, vaulting action and leaping action. Accordingly, the multiple preset areas include: a first area, a second area and a third area. The first area corresponds to the take-off action, that is, the area where the athlete is when performing the take-off action; the second area corresponds to the vaulting action, that is, the area where the athlete is when performing the vaulting action; the third area corresponds to the leaping action, that is, the area where the athlete is when performing the leaping action.

[0078] The multiple action states include: a first action state, a second action state, and a third action state. The first action state corresponds to a jumping action, which can be called a character jumping state; the second action state corresponds to a horse-supporting action, which can be called a character horse-supporting state; the third action state corresponds to a flying action, which can be called a character flying state.

[0079] In this embodiment, the above-mentioned target area, preset area and preset state can be applied to a specific scene, such as a video generation process of a horse vaulting scene, to improve the accuracy of video generation.

[0080] In some embodiments, the plurality of preset areas include: a first area, a second area, and a third area;

[0081] The preset state further includes: a first action state corresponding to the first area, a second action state corresponding to the second area, and a third action state corresponding to the third area;

[0082] The step of using the position relationship information as a current transition condition, starting from the first state being the no-action state, and sequentially determining a target state of the first state under each current transition condition as the second state includes:

[0083] If the first state is the no-action state, and the positional relationship information includes: the target area at least intersects with the first area, determining that the second state is the first action state; or,

[0084] If the first state is the first action state, and the position relationship information includes: the target area at least intersects with the second area, determining that the second state is the second action state; or,

[0085] If the first state is the second action state, and the position relationship information includes: the target area at least intersects with the third area, it is determined that the second state is the third action state; or,

[0086] If the first state is the third action state, and the position relationship information includes: the target area intersects with at least the third area or the second area, it is determined that the second state is the start state.

[0087] Furthermore, the target area at least intersects with the first area may include: the target area intersects with the first area, or the target area intersects with the first area and the second area; or

[0088] Furthermore, the target area at least intersects with the second area may include: the target area intersects with the second area, or the target area intersects with the second area and the third area; or,

[0089] Furthermore, the target area intersects with at least the third area, which may include: the target area intersects with the third area, the target area intersects with the first area and the third area, or the target area intersects with the first area, the second area and the third area; or,

[0090] Furthermore, the above-mentioned target area intersects with at least the third area or the second area, which may include: the target area intersects with the second area, the target area intersects with the second area and the third area, the target area intersects with the third area, the target area intersects with the first area and the third area, or the target area intersects with the first area, the second area and the third area.

[0091] refer to Figure 4 or Figure 5 The preset areas include the first area, the second area and the third area, which are represented by preset area A, preset area B and preset area C respectively; the target area is represented by area D. The intersection of the target area and the preset area can be referred to Figure 5 For example, the target area D intersects with the preset area A, the target area D intersects with the preset area A and the preset area B, the target area D intersects with the preset area B and the preset area C, and so on.

[0092] The default state machine is as follows Figure 6 As shown, reference Figure 6 , the circle represents the preset state of the preset state machine. Taking the horse vaulting scene as an example, the first action state is the character taking off, the second action state is the character holding the horse, and the third action state is the character flying; the images corresponding to each action state are as follows Figure 7As shown in the figure, the edges represent transition conditions. Transition condition A indicates that the target area intersects with preset area A, and transition condition A+B indicates that the target area intersects with both preset area A and preset area B. A and A+B are in an "or" relationship. The rest are similar and will not be repeated here.

[0093] Combine Figure 6 After the character completes the following actions in sequence: no action - character takes off - character supports the horse - character takes off into the air, a trigger event is generated.

[0094] In this embodiment, through the above-mentioned positional relationship information between the target area and the three preset areas, as well as the transition process from the first state to the second state, the state transition operation of specific scenes, such as the vaulting horse scene, can be completed, trigger events can be accurately generated, and the accuracy of the target video can be improved.

[0095] 305 . In response to the triggering event, determine a target image among the candidate images, and generate a target video based on the target image.

[0096] Specifically, the time point of occurrence of the trigger event is taken as the first time point, and the time point after the first time point and separated from the first time point by a first preset time length is taken as the second time point; the candidate image within the second preset time length before the second time point is obtained as the target image; and the target video is generated based on the target image.

[0097] In this embodiment, by determining the second time point through the first time point and obtaining a candidate image of a preset duration between the second time points as the target image, the comprehensiveness and accuracy of the target image can be improved, thereby improving the accuracy of the target video.

[0098] For example, the cache memory caches candidate images of a length of m, the first time point is represented by t1, and the first preset time length is assumed to be m / 2, then the second time point t2=t1+m / 2, assuming that the second preset time length is m, then the candidate images cached within the length of m can be read from the cache memory at the second time point t2 as target images, and the target images are composed into a target video.

[0099] In this embodiment, trigger events are generated by state machine state transitions, and target videos are generated based on these trigger events. This improves the accuracy and efficiency of data acquisition, ensuring that every detail of an athlete's movements can be accurately captured and analyzed. This is not only crucial for optimizing athlete training and competition strategies, but also provides coaches with more detailed and reliable data support. Furthermore, automated data recording and real-time analysis reduce the user's reliance on professional knowledge, making it easy for even non-professionals to use, thereby broadening the product's application range. Furthermore, through real-time feedback and analysis, users can adjust their training methods in a timely manner, effectively improving training results, enhancing product performance, and enhancing user experience.

[0100] Figure 8 is a schematic diagram according to the third embodiment of the present disclosure. This embodiment provides a video generation device, such as Figure 8 As shown, the device 800 includes: a detection module 801, a triggering module 802 and a generating module 803.

[0101] The detection module 801 is used to perform target detection processing on the candidate image to determine the target area in the candidate image; the trigger module 802 is used to generate a trigger event based on the positional relationship information between the target area and the preset area; the generation module 803 is used to determine the target image in the candidate image in response to the trigger event, and generate a target video based on the target image.

[0102] In this embodiment, by identifying a target image from candidate images and generating a target video based on the target image, the target video can be automatically generated, thereby avoiding the inefficiency and low accuracy issues associated with manual observation. Furthermore, the target image is acquired based on a trigger event, which is generated based on the positional relationship between the target area and a preset area. This allows for accurate triggering, thereby accurately acquiring the target image and improving the accuracy of the target video.

[0103] In some embodiments, the trigger module 802 is further used to: use the positional relationship information between the target area and the preset area as the current transfer condition, and transfer from the first state of the preset state machine to the second state according to the current transfer condition; the preset state of the preset state machine includes: the start state; if the second state is the start state, the trigger event is generated.

[0104] In this embodiment, the trigger event is generated based on the state machine, which can improve the accuracy of the trigger event and the accuracy of generating the target video.

[0105] In some embodiments, there are multiple preset areas;

[0106] The position relationship information is used to indicate that the target area has an intersection relationship with at least one of the multiple preset areas;

[0107] The preset state also includes: no action state;

[0108] The trigger module 802 is further configured to:

[0109] The position relationship information is used as the current transition condition, starting from the first state being the no-action state, the target state of the first state under the current transition condition is determined in sequence as the second state.

[0110] In this embodiment, the first state starts from the no-action state, and the second state is determined in sequence according to the current transition condition, which can improve processing accuracy.

[0111] In some embodiments, the target area is a body area of ​​the moving subject when performing a target movement; the target movement includes a plurality of preset key actions performed in sequence;

[0112] Each of the plurality of preset areas corresponds to each of the plurality of preset key actions;

[0113] The preset state further includes: a plurality of action states, and each action state in the plurality of action states corresponds one-to-one to each of the preset key actions.

[0114] In this embodiment, the above-mentioned target area, preset area and preset state can be applied to a specific scene, such as a video generation process of a horse vaulting scene, to improve the accuracy of video generation.

[0115] In some embodiments, the plurality of preset areas include: a first area, a second area, and a third area; the preset state further includes: a first action state corresponding to the first area, a second action state corresponding to the second area, and a third action state corresponding to the third area;

[0116] The trigger module 802 is further configured to:

[0117] If the first state is the no-action state, and the positional relationship information includes: the target area at least intersects with the first area, determining that the second state is the first action state; or,

[0118] If the first state is the first action state, and the position relationship information includes: the target area at least intersects with the second area, determining that the second state is the second action state; or,

[0119] If the first state is the second action state, and the position relationship information includes: the target area at least intersects with the third area, it is determined that the second state is the third action state; or,

[0120] If the first state is the third action state, and the position relationship information includes: the target area intersects with at least the third area or the second area, it is determined that the second state is the start state.

[0121] In this embodiment, through the above-mentioned positional relationship information between the target area and the three preset areas, as well as the transition process from the first state to the second state, the state transition operation of specific scenes, such as the vaulting horse scene, can be completed, trigger events can be accurately generated, and the accuracy of the target video can be improved.

[0122] In some embodiments, the generating module 803 is further configured to:

[0123] The time point at which the trigger event occurs is taken as a first time point, and a time point that is after the first time point and separated from the first time point by a first preset time length is taken as a second time point;

[0124] Acquire a candidate image within a second preset time period before the second time point as a target image;

[0125] A target video is generated based on the target image.

[0126] In this embodiment, by determining the second time point through the first time point and obtaining a candidate image of a preset duration between the second time points as the target image, the comprehensiveness and accuracy of the target image can be improved, thereby improving the accuracy of the target video.

[0127] In some embodiments, the detection module 801 is further configured to:

[0128] Determine a region of interest in the candidate image, and obtain a region of interest image corresponding to the region of interest;

[0129] Performing target detection processing on the region of interest image to determine the target region.

[0130] In this embodiment, target detection processing is performed on the region of interest. Since the size of the image of the region of interest is usually smaller than the candidate image, this can improve processing efficiency.

[0131] In some embodiments, the detection module 801 is further configured to:

[0132] The region of interest is determined based on the preset coordinates of the preset area.

[0133] In this embodiment, the region of interest is determined based on the preset coordinates of the preset area, so that a precise region of interest can be obtained, thereby improving processing accuracy.

[0134] It can be understood that in the embodiments of the present disclosure, the same or similar contents in different embodiments can be referenced to each other.

[0135] It can be understood that the terms “first”, “second”, etc. in the embodiments of the present disclosure are only used for distinction and do not indicate the degree of importance, time sequence, etc.

[0136] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0137] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0138] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device 900 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0139] like Figure 9 As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 909 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0140] Multiple components in the electronic device 900 are connected to the I / O interface 905, including an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0141] The computing unit 901 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as the video generation method. For example, in some embodiments, the video generation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the video generation method described above can be performed. Alternatively, in other embodiments, the computing unit 901 can be configured to perform the video generation method by any other appropriate means (e.g., by means of firmware).

[0142] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0143] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0144] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0145] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0146] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0147] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.

[0148] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0149] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A video generation method, comprising: Performing target detection processing on the candidate image to determine the target area in the candidate image; generating a trigger event based on the positional relationship information between the target area and the preset area; The trigger event is generated when the state of the state machine after the state transfer is the start state, and the state machine performs a state transfer operation based on the position relationship information; In response to the trigger event, determining a target image among the candidate images, and generating a target video based on the target image, comprising: The time point at which the trigger event occurs is taken as a first time point, and a time point that is after the first time point and separated from the first time point by a first preset time length is taken as a second time point; Acquire a candidate image within a second preset time period before the second time point as a target image; generating a target video based on the target image; The position relationship information is used to indicate that the target area has an intersection relationship with at least one preset area among the plurality of preset areas; The target area is the body area of ​​the moving object when performing the target movement; the target movement includes a plurality of preset key actions performed in sequence; Each of the plurality of preset areas corresponds to each of the plurality of preset key actions; The preset states of the state machine include: the startup state, the no-action state as the initial state, and multiple action states, and each action state in the multiple action states corresponds to each preset key action one by one; The trigger event is generated after the moving object completes the multiple preset key actions in sequence.

2. The method according to claim 1, wherein The generating of a trigger event based on the positional relationship information between the target area and the preset area includes: The positional relationship information between the target area and the preset area is used as the current transfer condition, and the preset state machine is transferred from the first state to the second state according to the current transfer condition; if the second state is the start state, the trigger event is generated.

3. The method according to claim 2, wherein: The method of using the positional relationship information between the target area and the preset area as a current transition condition and transitioning from the first state of the preset state machine to the second state according to the current transition condition includes: The position relationship information is used as the current transition condition, starting from the first state being the no-action state, the target state of the first state under the current transition condition is sequentially determined as the second state.

4. The method according to claim 3, wherein: The plurality of preset areas include: a first area, a second area and a third area; The plurality of action states include: a first action state corresponding to the first area, a second action state corresponding to the second area, and a third action state corresponding to the third area; The step of using the position relationship information as the current transition condition, starting from the first state being the no-action state, and sequentially determining the target state of the first state under the current transition condition as the second state includes: If the first state is the no-action state, and the positional relationship information includes: the target area at least intersects with the first area, determining that the second state is the first action state; or, If the first state is the first action state, and the position relationship information includes: the target area at least intersects with the second area, determining that the second state is the second action state; or, If the first state is the second action state, and the position relationship information includes: the target area at least intersects with the third area, it is determined that the second state is the third action state; or, If the first state is the third action state, and the position relationship information includes: the target area intersects with at least the third area or the second area, it is determined that the second state is the start state.

5. The method according to any one of claims 1 to 4, wherein: The performing target detection processing on the candidate image to determine the target area in the candidate image includes: Determine a region of interest in the candidate image, and obtain a region of interest image corresponding to the region of interest; Performing target detection processing on the region of interest image to determine the target region.

6. The method according to claim 5, wherein: The determining of the region of interest in the candidate image includes: The region of interest is determined based on the preset coordinates of the preset area.

7. A video generation device, comprising: A detection module, configured to perform target detection processing on the candidate image to determine a target area in the candidate image; A trigger module, configured to generate a trigger event based on positional relationship information between the target area and the preset area; The trigger event is generated when the state of the state machine after the state transfer is the start state, and the state machine performs a state transfer operation based on the position relationship information; a generating module, configured to determine a target image among the candidate images in response to the triggering event, and generate a target video based on the target image; The generating module is further configured to: The time point at which the trigger event occurs is taken as a first time point, and a time point that is after the first time point and separated from the first time point by a first preset time length is taken as a second time point; Acquire a candidate image within a second preset time period before the second time point as a target image; generating a target video based on the target image; The position relationship information is used to indicate that the target area has an intersection relationship with at least one preset area among the plurality of preset areas; The target area is the human body area of ​​the moving object when performing the target movement; The target movement includes a plurality of preset key actions to be performed in sequence; Each of the plurality of preset areas corresponds to each of the plurality of preset key actions; The preset states of the state machine include: the startup state, the no-action state as the initial state, and multiple action states, and each action state in the multiple action states corresponds to each preset key action one by one; The trigger event is generated after the moving object completes the multiple preset key actions in sequence.

8. The device according to claim 7, wherein The trigger module is further configured to: Using the positional relationship information between the target area and the preset area as a current transition condition, and transitioning from the first state of the preset state machine to the second state according to the current transition condition; If the second state is the start state, the trigger event is generated.

9. The device according to claim 8, wherein The trigger module is further configured to: The position relationship information is used as the current transition condition, starting from the first state being the no-action state, the target state of the first state under the current transition condition is determined in sequence as the second state.

10. The device according to claim 9, wherein The plurality of preset areas include: a first area, a second area and a third area; The plurality of action states include: a first action state corresponding to the first area, a second action state corresponding to the second area, and a third action state corresponding to the third area; The trigger module is further configured to: If the first state is the no-action state, and the positional relationship information includes: the target area at least intersects with the first area, determining that the second state is the first action state; or, If the first state is the first action state, and the position relationship information includes: the target area at least intersects with the second area, determining that the second state is the second action state; or, If the first state is the second action state, and the position relationship information includes: the target area at least intersects with the third area, it is determined that the second state is the third action state; or, If the first state is the third action state, and the position relationship information includes: the target area intersects with at least the third area or the second area, it is determined that the second state is the start state.

11. The device according to any one of claims 7 to 10, wherein: The detection module is further configured to: Determine a region of interest in the candidate image, and obtain a region of interest image corresponding to the region of interest; Performing target detection processing on the region of interest image to determine the target region.

12. The device according to claim 11, wherein The detection module is further configured to: The region of interest is determined based on the preset coordinates of the preset area.

13. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

14. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.

15. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Display method and device in augmented reality scene, electronic equipment and storage medium

    CN112653848A

  • Video processing method and device, electronic equipment and medium

    CN115022679A