Method, device, system, medium and equipment for recognizing target object in video

By acquiring feature points that reflect three-dimensional spatial information in videos, determining the detection area, and calculating the number of repetitions, the problem of inaccurate target object recognition in videos is solved, and efficient and accurate target object counting is achieved.

CN114821460BActive Publication Date: 2026-05-15ZHEJIANG E COMMERCE BANK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG E COMMERCE BANK CO LTD
Filing Date
2022-03-17
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify the number of target objects in videos, resulting in low accuracy in scenarios such as asset inventory.

Method used

By acquiring an image sequence containing multiple images, each image corresponding to feature points reflecting three-dimensional spatial information, the detection area of ​​the target object in adjacent images is determined, and the number of repetitions is calculated based on the comparison of feature points in the detection area to identify the number of target objects.

Benefits of technology

It enables accurate counting of target objects in videos, improves recognition accuracy and efficiency, and enhances recognition performance in scenarios such as asset inventory.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821460B_ABST
    Figure CN114821460B_ABST
Patent Text Reader

Abstract

The embodiment of the present specification provides a kind of identification method, device, system, computer readable storage medium and electronic equipment of target object in video, which comprises: obtaining the image sequence containing multiple images from target video, wherein each image corresponds to feature point reflecting three-dimensional space information, so that the three-dimensional coordinate information of target object can be obtained according to the video shot.For two adjacent images in image sequence, first, the detection area where each target object is located in two images is determined respectively.Then, the feature point corresponding to each detection area is obtained.As the feature point corresponding to detection area reflects the three-dimensional space information of target object, further, the number of repeatedly shot target objects in two adjacent images can be determined based on the comparison between the feature points corresponding to detection area between two images, so that the identification of the number of target objects in target video can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of video processing technology, and in particular to a method, apparatus, system, computer-readable storage medium, and electronic device for identifying target objects in a video. Background Technology

[0002] By recording videos, users can remotely observe the objects depicted in the videos. For example, if user A records a video of a herd of pigs, user B can remotely monitor the situation of the pigs. However, it's difficult to accurately identify the actual condition of the objects in a video solely based on the footage; for instance, it's impossible to accurately determine the number of pigs in the video.

[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this specification, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention

[0004] The purpose of this specification is to provide a method, apparatus, system, computer-readable storage medium, and electronic device for identifying target objects in videos, thereby improving the accuracy and efficiency of object identification in videos to at least a certain extent.

[0005] Other features and advantages of this specification will become apparent from the following detailed description, or may be learned in part by practice of this specification.

[0006] According to one aspect of this specification, a method for identifying a target object in a video is provided, applied to a server. The method includes: acquiring an image sequence containing L images from a target video, wherein each image in the image sequence corresponds to feature points reflecting three-dimensional spatial information, and L is an integer greater than 1; determining a detection region corresponding to a target object in the s-th image and determining a detection region corresponding to the target object in the (s+1)-th image, each detection region corresponding to one target object, and s taking values ​​from 1 to L-1; and determining the number X of repetitions of the target object between the s-th image and the (s+1)-th image based on the feature points corresponding to the detection regions in the s-th image and the detection regions in the (s+1)-th image. s,s+1 ; and, based on the above repetition count X s,s+1 Identify the number of the aforementioned target objects in the target video.

[0007] According to another aspect of this specification, a method for identifying a target object in a video is provided, applied to a terminal. The method includes: capturing a target video; and sending the target video to a server, so that the server: and obtains an image sequence containing L images from the target video, wherein each image in the image sequence corresponds to feature points reflecting three-dimensional spatial information, and L is an integer greater than 1; determining a detection region corresponding to the target object in the s-th image and determining a detection region corresponding to the target object in the (s+1)-th image, each detection region corresponding to one target object, and s taking the value of an integer from 1 to L-1; and determining the number X of repetitions of the target object between the s-th image and the (s+1)-th image based on the feature points corresponding to the detection regions in the s-th image and the detection regions in the (s+1)-th image. s,s+1 ; and, based on the above repetition count X s,s+1 Identify the number of the aforementioned target objects in the target video.

[0008] According to another aspect of this specification, a device for identifying target objects in a video is provided, configured on a server, the device comprising: a feature point determination module, a detection region determination module, a repetition count determination module, and an identification module.

[0009] The feature point determination module is used to acquire an image sequence containing L images from the target video, wherein each image in the image sequence corresponds to a feature point reflecting three-dimensional spatial information, and L is an integer greater than 1; the detection region determination module is used to determine the detection region corresponding to the target object in the s-th image and the detection region corresponding to the target object in the (s+1)-th image, wherein each detection region corresponds to one target object, and s takes the value of an integer from 1 to L-1; the repetition number determination module is used to determine the repetition number X of the target object between the s-th image and the (s+1)-th image based on the feature points corresponding to the detection regions in the s-th image and the detection regions in the (s+1)-th image. s,s+1 ; and the aforementioned identification module is used to identify the number of repetitions X. s,s+1 Identify the number of the aforementioned target objects in the target video.

[0010] According to another aspect of this specification, a device for identifying target objects in a video is provided, configured in a terminal, the device comprising: a video capturing module and a transmitting module.

[0011] The video capturing module is used to capture the target video; and the sending module is used to send the target video to the server, so that the server:

[0012] From the target video, obtain an image sequence containing L images, where each image in the image sequence corresponds to feature points reflecting three-dimensional spatial information, and L is an integer greater than 1; determine the detection region corresponding to the target object in the s-th image and the detection region corresponding to the target object in the (s+1)-th image, where each detection region corresponds to one target object, and s takes the value of an integer from 1 to L-1; based on the feature points corresponding to the detection regions in the s-th and (s+1)-th images, determine the number of repetitions X of the target object between the s-th and (s+1)-th images. s,s+1 ; and, based on the above repetition count X s,s+1 Identify the number of the aforementioned target objects in the target video.

[0013] According to one aspect of this specification, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for identifying a target object in a video as described in the above embodiments.

[0014] According to another aspect of this specification, a computer-readable storage medium is provided that stores instructions which, when executed on a computer or processor, cause the computer or processor to perform the method for identifying a target object in a video as described in the above embodiments.

[0015] The methods, apparatus, systems, computer-readable storage media, and electronic devices for identifying target objects in videos provided in the embodiments of this specification have the following technical effects:

[0016] In the exemplary embodiments provided in this specification, an image sequence containing multiple images is obtained from a target video. Each image corresponds to feature points reflecting three-dimensional spatial information, thereby allowing the three-dimensional coordinate information of the target object to be obtained from the captured video. Further, for two adjacent images in the image sequence, the detection region containing each target object is first determined in each image. For example, the first image contains M detection regions, and the second image contains N detection regions. Then, the feature points corresponding to each detection region are obtained. As mentioned above, since the feature points corresponding to the detection regions reflect the three-dimensional spatial information of the target object, the number of repeatedly captured target objects in two adjacent images can be determined by comparing the feature points corresponding to the detection regions between the two images. Therefore, this solution, based on the feature points reflecting three-dimensional spatial information carried in the captured video, can accurately count the target objects in the video, thereby effectively improving the accuracy and efficiency of object recognition in the video.

[0017] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description

[0018] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this specification and, together with the description, serve to explain the principles of this specification. It is obvious that the drawings described below are merely some embodiments of this specification, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0019] Figure 1 This is a schematic diagram illustrating the architecture of the target object recognition scheme in a video provided in the embodiments of this specification.

[0020] Figure 2 This is a flowchart illustrating a method for identifying target objects in a video, provided as an embodiment of this specification.

[0021] Figure 3 This is a schematic diagram illustrating the information interaction of a method for identifying target objects in a video, provided as another embodiment of this specification.

[0022] Figure 4a and Figure 4b This is a schematic diagram of adjacent images in an image sequence provided in one embodiment of this specification.

[0023] Figure 5 This is a schematic diagram of a video shooting scene for a target object provided in one embodiment of this specification.

[0024] Figure 6This is a flowchart illustrating a method for determining the number of repetitions of a target object between adjacent images in an image series, as provided in an embodiment of this specification.

[0025] Figure 7 This is a flowchart illustrating a method for determining the number of repetitions of a target object between adjacent images in an image series, provided as another embodiment of this specification.

[0026] Figure 8 This is a flowchart illustrating a method for determining the number of repetitions of a target object between adjacent images in an image series, as provided in another embodiment of this specification.

[0027] Figure 9 This is a schematic diagram of a video shooting scene for a target object, provided as another embodiment of this specification.

[0028] Figure 10 This is a flowchart illustrating a method for identifying target objects in a video, provided as another embodiment of this specification.

[0029] Figure 11 This is a schematic diagram of the structure of a device for recognizing target objects in a video, provided in one embodiment of this specification.

[0030] Figure 12 This is a schematic diagram of the structure of a device for recognizing target objects in a video, provided as another embodiment of this specification.

[0031] Figure 13 This is a schematic diagram of the structure of a device for recognizing target objects in a video, provided in another embodiment of this specification.

[0032] Figure 14 This is a schematic diagram of the structure of a device for recognizing target objects in a video, provided as another embodiment of this specification.

[0033] Figure 15 This is a schematic diagram of the structure of a target object recognition system in a video provided in one embodiment of this specification.

[0034] Figure 16 This is a schematic diagram of the structure of a target object recognition system in a video provided for another embodiment of this specification.

[0035] Figure 17 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this specification. Detailed Implementation

[0036] To make the objectives, technical solutions, and advantages of this specification clearer, the embodiments of this specification will be described in further detail below with reference to the accompanying drawings.

[0037] In the following description, when referring to the accompanying drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.

[0038] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided to make this specification more comprehensive and complete, and to fully convey the concept of example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a full understanding of the embodiments described herein. However, those skilled in the art will recognize that the technical solutions described herein may be practiced with one or more of the specific details omitted, or other methods, components, apparatus, steps, etc., may be employed. In other instances, well-known technical solutions are not shown or described in detail to avoid obscuring various aspects of this specification.

[0039] Furthermore, the accompanying drawings are merely illustrative diagrams of this specification and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0040] One application scenario for identifying target objects in videos, as provided in the embodiments of this specification, is an asset inventory scenario. Specifically, if a user needs to provide asset proof to MYbank to demonstrate their asset status, the user needs to film their assets, such as farms, real estate, and vehicles. Therefore, MYbank needs to verify the authenticity and quantity of the assets (objects) filmed in the relevant videos to grant appropriate credit to the user based on the asset inventory results.

[0041] Related technologies employ image recognition algorithms to identify images in videos. However, because the images captured by these technologies only reflect the corresponding two-dimensional information, there's a risk of duplicate captures of the same object in adjacent video frames (e.g., a group of pigs 'a' may appear in an earlier frame and a portion of 'a' may appear in a later frame). This can lead to duplicate counting of the same asset, resulting in inaccurate identification. Therefore, for asset inventory scenarios, these technologies suffer from low accuracy in overall asset counting.

[0042] The embodiments in this specification solve the above-mentioned technical problems and provide this technical solution. Specifically, Figure 1 This is a schematic diagram of the system architecture for the target object recognition scheme in the video provided in the embodiments of this specification.

[0043] like Figure 1 As shown, the system architecture may include a terminal 110, a network 130, and a server 120. The terminal 110 and the server 120 are connected via the network 130.

[0044] Terminal 110 can be a mobile phone, computer, tablet, etc., containing a camera component, or a camera with video recording capabilities. Network 130 can be a communication medium of various connection types that can provide a communication link between terminal 110 and server 120, such as a wired communication link, a wireless communication link, or a fiber optic cable, etc., which are not limited herein. Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, and big data and artificial intelligence platforms.

[0045] In the embodiments provided in this specification, terminal 110 captures target video of assets related to the breeding farm 11. It should be noted that in this embodiment, the captured video can acquire three-dimensional spatial data of the captured object. For example, by using a software development kit (SDK) based on the Simultaneous Localization and Mapping (SLAM) framework installed on the terminal, the captured video images can correspond to feature points reflecting three-dimensional spatial information. Furthermore, terminal 110 needs to be equipped with an inertial measurement unit (IMU) sensor, gyroscope, geomagnetic sensor, etc., to facilitate the server in determining the video capture trajectory and direction based on the video.

[0046] Furthermore, terminal 110 sends the target video to server 120 via network 130.

[0047] On one hand, server 120 acquires an image sequence containing multiple images from the target video, where each image corresponds to feature points reflecting three-dimensional spatial information, thereby obtaining the three-dimensional coordinate information of the target object based on the captured video. Further, for two adjacent images in the image sequence, server 120 first determines the detection region where each target object is located in each image; for example, the first image contains M detection regions, and the second image contains N detection regions. Then, it acquires the feature points corresponding to each detection region. As mentioned above, since the feature points corresponding to the detection regions reflect the three-dimensional spatial information of the target object, further, based on the comparison between the feature points corresponding to the detection regions in two images, the number of target objects repeatedly captured in the two adjacent images can be determined, thus obtaining the target object recognition result 12. It can be seen that this solution, based on the feature points reflecting three-dimensional spatial information carried in the captured video, can achieve accurate counting of target objects in the video, thereby effectively improving the accuracy and efficiency of object recognition in the video.

[0048] On the other hand, the server 120 obtains the shooting trajectory and shooting direction of the terminal based on the target video, and further generates shooting suggestions 13. The server 120 sends the above shooting suggestions 13 to the terminal 110 in real time through the network 130. Thus, the user can refer to the shooting suggestions 13 during the shooting process, which helps to improve the quality of the target video and facilitates the identification of the target object in the target video.

[0049] The following is passed first Figures 2 to 10 This specification provides a detailed description of the embodiments of the method for identifying target objects in videos:

[0050] For example, Figure 2 This is a flowchart illustrating a method for identifying target objects in a video, as provided in an embodiment of this specification. (Reference) Figure 2 The method shown in this embodiment includes: S210-S240.

[0051] In S210, an image sequence containing L images is obtained from the target video. Each image in the image sequence corresponds to a feature point that reflects three-dimensional spatial information, and L is an integer greater than 1.

[0052] For example, refer to Figure 1In the asset inventory scenario of farm 11, the aforementioned target video is obtained based on Structure From Motion (SFM) technology or the SLAM framework, and by capturing images of relevant target objects using various sensors and cameras. For example, the user can obtain a complete view of the farm by launching the SLAM SDK installed on the terminal and moving the terminal to capture images of the target objects. It should be noted that the captured video is not limited to SFM technology and the SLAM framework; it can also be other technologies capable of acquiring three-dimensional spatial information reflecting the target objects. This specification does not limit the scope of the embodiments described herein.

[0053] For example, the server extracts images from the target video to obtain a sequence of images in chronological order. Since the target video was captured using motion structure reconstruction technology or simultaneous localization and framing, feature points reflecting three-dimensional spatial information can be obtained from each image in the sequence.

[0054] In S220, the detection region corresponding to the target object in the s-th image of the image sequence and the detection region corresponding to the target object in the (s+1)-th image are determined. Each detection region corresponds to a target object, and the value of s is an integer from 1 to L-1.

[0055] For example, computer recognition algorithms can be used to determine bounding boxes in each image of an image sequence, with each bounding box corresponding to a target object. For instance: Figure 4a For the s-th image, Figure 4b For the (s+1)th image, multiple detection boxes 41 are identified in both the s-th and (s+1)-th images, with each detection box 41 containing a detection region 42 corresponding to a pig. Alternatively, detection regions in each image can be identified using artificial intelligence algorithms. It should be noted that each detection region corresponds to a target object, and the shape of the detection region is not limited to a square. For example, the shape of the detection region can also be an irregular shape corresponding to the shape of the target object (such as a pig) as presented in the current image.

[0056] For example, because video images with adjacent timestamps are prone to being captured repeatedly, for example, Figure 4a The dashed detection box in (image s) and Figure 4b The dashed detection box in the (s+1)th image represents repeated images of the same pig herd. Therefore, to address the issue of counting repetitions of the target object, we focus on two adjacent images in the image sequence. In this embodiment, for two adjacent images in the image sequence: the s-th image and the (s+1)-th image, we determine the detection region for each image.

[0057] In S230, based on the feature points corresponding to the detection region in the s-th image and the feature points corresponding to the detection region in the (s+1)-th image, the number of repetitions X of the target object between the s-th image and the (s+1)-th image is determined. s,s+1 .

[0058] In this embodiment, since the feature points corresponding to the detection area reflect the three-dimensional spatial information of the target object, the number of target objects repeatedly captured in two adjacent images can be determined by comparing the feature points corresponding to the detection areas between two adjacent images. That is, the number of repetitions of the target object between the s-th image and the (s+1)-th image above is X. s,s+1 .

[0059] For example, in SLAM, feature points are typically computationally less computationally intensive Oriented Fast and Rotated Brief (ORB) feature points or feature points exhibiting optical flow variations. In SFM, they are typically more computationally intensive but more stable feature points, such as those derived from Scale Invariant Feature Transform (SIFT). Each feature point corresponds to a specific 3D coordinate system (x, y, z), where the origin can be specified by the algorithm, for example, the starting point of the image capture.

[0060] For example, for each image, feature points corresponding to the identified detection regions in that image can be obtained, thus obtaining feature points for that image used to calculate the number of repetitions of the target object with neighboring images. Alternatively, all feature points in each image can be clustered, discarding feature points that are far from each cluster center, and using the remaining feature points as feature points for calculating the number of repetitions of the target object with neighboring images.

[0061] In S240, based on the number of repetitions X s,s+1 Identify the number of target objects in the target video.

[0062] For example, in S220, the number of detection regions identified in each image corresponds to the total number of target objects presented in that image. For instance, if the s-th image contains M detection regions, then the total number of target objects presented in the s-th image is M; if the (s+1)-th image contains N detection regions, then the total number of target objects presented in the (s+1)-th image is N. Further, the sum of M and N is calculated, and then the sum is multiplied by X. s,s+1 The difference is determined as the actual number of target objects in the s-th image and the (s+1)-th image.

[0063] When s takes the value of an integer from 1 to L-1, the actual number of target objects in the target video can be obtained. For example, if s is 1, then according to the above embodiment, X can be obtained. 1,2 If s takes the value 2, then according to the above embodiment, X can be obtained. 2,3 ... s takes the value L-1, then according to the above embodiment, X can be obtained L-1,L .

[0064] Furthermore, if through number s Let represent the number of target objects contained in the s-th image. Then, the actual number of target objects in the L images of the above image sequence can be obtained as: number1 + number2 + ... + number L-1 +number L -(X 1,2 +X 2,3 +……+X L-1,L ).

[0065] This instruction manual Figure 2 In the solution provided by the embodiment shown, based on the feature points reflecting three-dimensional spatial information carried in the captured video, the target objects in the video can be accurately counted, thereby effectively improving the accuracy of object recognition in the video and also improving the efficiency of object recognition in the video.

[0066] In an exemplary embodiment, Figure 3 This diagram illustrates the information interaction of a method for identifying target objects in a video, as provided in another embodiment of this specification. Specifically, it describes the implementation process of the target object identification scheme provided in this embodiment from the perspective of information interaction between terminal 110 and server 120.

[0067] refer to Figure 3 In S310, terminal 110 captures target video based on motion structure recovery technology or simultaneous positioning and mapping frame, as well as based on multiple sensors and cameras to capture the target object.

[0068] For example, when terminal 110 uses sensing components such as IMU and magnetometer and runs the SLAM SDK, it can capture images such as... Figure 5 The above-mentioned target video was obtained from the shown breeding farm.

[0069] For example, the shooting direction of the terminal camera should be as perpendicular as possible to the tangent direction of the shooting trajectory, which is more conducive to obtaining more accurate three-dimensional information.

[0070] Furthermore, in S320, terminal 110 sends the target video to server 120. And, as a specific implementation of S210: in S330, server 120 extracts images from the target video to obtain an image sequence containing L images based on one or more of the following factors: the terminal's moving speed, the terminal's corresponding bandwidth, and the type of the target object.

[0071] Considering the efficiency and computational load of identifying target objects in the target video, in this embodiment, key images will be extracted from the target video based on the terminal's moving speed, the terminal's corresponding bandwidth, and the type of target object, to obtain an image sequence with a chronological order of shooting time.

[0072] For example, the slower the terminal's movement speed, the narrower the terminal's bandwidth, and the larger the target object's outline size during shooting, the longer the time interval for capturing images from the target video. Conversely, the faster the terminal's movement speed, the wider the terminal's bandwidth, and the smaller the target object's outline size during shooting, the shorter the time interval for capturing images from the target video. Therefore, the time interval for capturing images from the target video needs to be determined based on the actual situation; this article does not impose a limit on the duration of the time interval for capturing images from the target video.

[0073] In this embodiment, in the target video captured based on motion structure recovery technology or simultaneous positioning and framing, the server 120 can obtain feature points reflecting three-dimensional spatial information corresponding to each image based on the target video for the purpose of identifying the number of target objects. This will be explained through the embodiments corresponding to S210-S240. On the other hand, the server 120 can obtain the shooting trajectory and camera shooting direction based on the target video and use them to generate shooting suggestions, specifically:

[0074] In S340, server 120 determines the shooting trajectory of the target object and the shooting direction of the camera based on the target video; in S350, server 120 determines the tangent direction of the target shooting point based on the shooting trajectory, calculates the angle between the tangent direction of the target shooting point and the shooting direction of the camera, and determines the shooting suggestion for the target shooting point based on the angle.

[0075] refer to Figure 5 Server 120 can obtain the shooting trajectory 500 based on the target video. For example, the shooting trajectory 500 sequentially includes shooting points 51, 52, and 53. Furthermore, the shooting direction of the camera 50 at each shooting point can also be obtained based on the target video. (Reference) Figure 5It is evident that the field of view (FOV) varies depending on the shooting direction at each shooting point. To ensure the captured video yields more accurate 3D information, the camera's shooting direction should be as perpendicular as possible to the tangent of the shooting trajectory. For example, at shooting point 51, the shooting direction should be as perpendicular as possible to the tangent 51' of the shooting trajectory 500 at that point. Therefore, when the captured video is transmitted to server 120 in real time (terminal 110 simultaneously shoots and sends video to server 120), server 120 can determine shooting suggestions for the current shooting point in real time. For example, refer to... Figure 5 For the current shooting point 53, the server 120 determines the tangent direction of shooting point 53 based on the shooting trajectory 500, and calculates the angle between the tangent direction of shooting point 53 and the actual shooting direction of the camera. If the angle is not equal to 90 degrees, the difference in angle and direction between the actual shooting direction and the target shooting direction (which has an angle of 90 degrees with the tangent direction) is determined as the shooting suggestion for the current shooting point.

[0076] Furthermore, in S360, server 120 sends shooting suggestions to terminal 110 in real time; and in S370, terminal 110 displays the shooting suggestions. For example, augmented reality (AR) technology can be used to display shooting suggestions corresponding to each shooting point at the respective shooting point. For example, refer to... Figure 5 When a user is wearing augmented reality equipment, at the current shooting location (point 53), relevant shooting suggestions can be displayed at the user's location. These suggestions can be directional arrows, allowing the user to adjust the camera's shooting direction accordingly. By following the real-time shooting suggestions, the user can ensure that the camera's shooting direction at each shooting point is as perpendicular as possible to the tangent of the shooting trajectory, thus improving the quality of the target video.

[0077] pass Figure 3 The provided embodiments determine the shooting trajectory and camera shooting direction based on the target video, thereby providing shooting suggestions. Furthermore, guiding the user to shoot the scene at an accurate angle through AR interaction improves the shooting quality of the target video, which in turn facilitates accurate identification of target objects based on the video.

[0078] In an exemplary embodiment, as a specific implementation of S230, Figures 6 to 8 Flowcharts are provided for methods to determine the number of repetitions of a target object between adjacent images in an image series.

[0079] refer to Figure 6In S610, the s-th image contains M detection regions. The center point C is determined based on the feature points corresponding to the ith detection region in the s-th image. i M is a positive integer, and i takes any positive integer value from 1 to M. Also, in S620, the (s+1)th image contains N detection regions, and the center point C' is determined based on the feature points corresponding to the j-th detection region in the (s+1)th image. j N is a positive integer, and j takes the value of any positive integer from 1 to N.

[0080] For example, refer to Figure 4a The s-th image contains M detection regions, and the (s+1)-th image contains N detection regions. The center point will be determined based on the feature points corresponding to each detection region.

[0081] Taking the i-th detection region in the s-th image as an example, since each feature point can be represented as a three-dimensional coordinate in the same coordinate system, the feature points corresponding to the i-th detection region in the s-th image can determine a three-dimensional space V. i For example, the three-dimensional space V can be... i The center is taken as the above center point C i Specifically, the center point C can be determined. i 3D coordinates (x) is y is , z is Similarly, the feature point center C' corresponding to the j-th detection region in the (s+1)-th image can be determined. j 3D coordinates (x) j(s+1) y j(s+1) , z j(s+1) ).

[0082] As can be seen, in this embodiment, the center point corresponding to each detection area may be a feature point or may not be an original feature point, but both can be represented by three-dimensional coordinates in the same coordinate system.

[0083] Continue to refer to Figure 6 In S630, the center point C is calculated. i With center point C' j Distance D between i,j And based on distance D i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 .

[0084] For example, based on center point C i 3D coordinates (x) is y is , z is ) and center point C' j 3D coordinates (x)j(s+1) y j(s+1) , z j(s+1) The distance D can be calculated. i,j ,like:

[0085]

[0086] For example, the first preset value mentioned above is related to the type of the target object; for instance, it can be the average of the maximum outer dimensions of the target object. In the case of a pig, the first preset value is, for example, 1.2 meters. If the distance D is calculated... i,j If the distance is greater than 1.2 meters, it means that the target object corresponding to the i-th detection region in the s-th image and the target object corresponding to the j-th detection region in the (s+1)-th image are not the same pig; if the distance D is calculated... i,j If the distance is no greater than 1.2 meters, it means that the target object corresponding to the i-th detection region in the s-th image and the target object corresponding to the j-th detection region in the (s+1)-th image belong to the same pig.

[0087] For example, in the process of calculating the distance between each center point in adjacent images, each time the distance D occurs... i,j If the number of repetitions is not greater than the first preset value, that is, if there is a case where the target object corresponding to the i-th detection region in the s-th image and the target object corresponding to the j-th detection region in the (s+1)-th image belong to the same target object, the repetition count is increased by one. This allows us to determine the total repetition count X between adjacent s-th and (s+1)-th images. s,s +1.

[0088] In an exemplary embodiment, Figure 7 The illustrated embodiment can be considered as another specific implementation of S230.

[0089] refer to Figure 7 In S710, the feature points corresponding to the detection region in the s-th image are clustered to obtain M cluster centers corresponding to the s-th image, where M is a positive integer; and in S720, the feature points corresponding to the detection region in the (s+1)-th image are clustered to obtain N cluster centers corresponding to the (s+1)-th image, where N is a positive integer.

[0090] For example, respectively for Figure 4a (Image s) and Figure 4b In the (s+1)th image, computer vision (CV) recognition of pigs can be performed, resulting in detection boxes 41 for each pig in the s-th image and 41 for each pig in the (s+1)-th image. Correspondingly, the number of pigs in the s-th image can be counted. s=M, the number of pigs in the (s+1)th image. s+1 =N.

[0091] Furthermore, taking the s-th image as an example, for obtaining all feature points 43 within the detection boxes in the s-th image, three-dimensional spatial clustering is performed on the feature points (e.g., k-means algorithm or DBSCAN algorithm). For example, the cluster centers in the s-th image are recorded as c1. s c2 s c3 s Similarly, the cluster centers in the (s+1)th image can be obtained: c1 s+1 c2 s+1 c3 s+1 ...

[0092] In S730, the i-th cluster center C in the s-th image is calculated. i The j-th cluster center C' in the (s+1)-th image j Distance D between i,j Let i take the values ​​of all positive integers from 1 to M, and j take the values ​​of all positive integers from 1 to N; and in S740, according to the distance D i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 .

[0093] For example, calculate the pairwise distance D between the cluster centers in the (s+1)th image and the cluster centers in the sth image. i,j =range(ci) s cj s+1 Furthermore, compare D. i,j =range((ci s cj s+1 The magnitude of D compared to the aforementioned first preset value. As mentioned before, in D i,j =range(ci) s cj s+1 If the value is less than the first preset value mentioned above, it indicates that the pig was photographed repeatedly in the two images. For example, D 5,3 =range((c5) s c3 s+1 If the value is less than the first preset value mentioned above, then the pig corresponding to the 5th cluster center in the s-th image and the pig corresponding to the 3rd cluster center in the (s+1)-th image are considered duplicate photos. The number of duplicates X of the target object can then be calculated. s,s+1 .

[0094] Furthermore, the actual number of target objects in the L images of the above image sequence can be calculated as: number1 + number2 + ... + numberL-1 +number L -(X 1,2 +X 2,3 +X L-1,L ).

[0095] In an exemplary embodiment, Figure 8 The illustrated embodiment can be considered as another specific implementation of S230.

[0096] refer to Figure 8 In S810, the s-th image contains M detection regions, and the sparse point cloud D corresponding to the i-th detection region in the s-th image is obtained. i Given that the (s+1)th image contains N detection regions, obtain the sparse point cloud D' corresponding to the j-th detection region in the (s+1)th image. j .

[0097] For the same target object, due to shooting angle or occlusion by other objects, it is impossible to obtain the 3D point cloud of the entire object in a single scan. Therefore, in this embodiment, feature points of the target object in multiple images are matched in 3D to obtain a 3D graphic (point cloud) after stitching together based on the feature points.

[0098] For example, the i-th detection region contained in the first s images of the image sequence is obtained, resulting in s detection regions. The feature points corresponding to the s detection regions are then matched to obtain the sparse point cloud D corresponding to the i-th detection region in the s-th image. i .

[0099] For example, the 3D matching algorithm used can be the Iterative ClosestPoint (ICP) algorithm or the global matching algorithm, but it is not limited to ICP and the global matching algorithm. Other algorithms can also be used for 3D matching.

[0100] For example, the j-th detection region contained in the first s+1 images of the image sequence is obtained to get s+1 detection regions, and the feature points corresponding to the s+1 detection regions are matched to obtain the sparse point cloud D' corresponding to the j-th detection region in the s+1-th image. j .

[0101] In this embodiment, since feature points can reflect the three-dimensional spatial information of the target object, point clouds can be obtained through feature point matching. Point clouds can accurately and vividly reflect the current spatial geometry of the target object corresponding to each detection area, which is beneficial for accurately determining the repetition of the target object between adjacent images.

[0102] Continue to refer to Figure 8In S820, the sparse point cloud D is determined. i Corresponding center point C i And determine the sparse point cloud D' j The corresponding center point C' j j takes the value of any positive integer from 1 to N; and, in S830, the center point C is calculated. i With center point C' j Distance D between i,j And based on distance D i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 .

[0103] For example, specific implementations of S820 and S830 can be found in [reference]. Figure 6 The specific implementation methods corresponding to the embodiments shown will not be described in detail here.

[0104] In an exemplary embodiment, during the asset inventory process, in addition to identifying the quantity of assets, it is also necessary to verify the authenticity of the photographed assets. For example, Figure 9 This is a schematic diagram of a video shooting scene for a target object, provided as another embodiment of this specification.

[0105] refer to Figure 9 The toy truck 90 in the video could potentially be used as a vehicle in the filming. To verify the authenticity of the target object in the target video, embodiments of this specification provide methods such as... Figure 10 The solution.

[0106] refer to Figure 10 In S810, obtain such Figure 9 Sparse point cloud D of the target object "truck" i The specific implementation of S810 will not be described in detail. Further, in S840, the sparse point cloud D is acquired. i The maximum envelope size is determined; the maximum envelope size is compared with the target envelope size, and the identification result regarding the authenticity of the target object is determined based on the comparison result. The target envelope size is determined according to the category of the target object.

[0107] For example, the maximum envelope size mentioned above can be the maximum size in the minimum envelope map. For instance, the maximum size in the minimum envelope map T1 of the sparse point cloud in the cross section and the minimum envelope map T2 in the longitudinal section is determined in the minimum envelope map T1 and the minimum envelope map T2, and this maximum size is determined as the maximum envelope size mentioned above.

[0108] For example, the sparse point cloud D of a "truck" iThe maximum envelope size is 0.5 meters. The target envelope size for the "truck" is 15 meters. Comparing the "truck's" maximum envelope size of 0.5 meters with the target envelope size of 15 meters, it is clear that the "truck's" maximum envelope size of 0.5 meters is much smaller than the target envelope size of 15 meters. Therefore, it can be determined that... Figure 9 The "truck" shown in the image is not a real truck, thus confirming the authenticity of the target object.

[0109] The target object recognition scheme in the video provided in this specification's embodiments includes feature points of the captured object in the video provided by the terminal. These feature points reflect the object's three-dimensional spatial information (such as three-dimensional coordinates). On one hand, based on the feature points corresponding to adjacent images in the image sequence, the repetition count of the target object between adjacent images can be determined, obtaining the actual quantity of the target object, for example, enabling asset inventory. Simultaneously, based on the feature points reflecting the object's three-dimensional spatial information, the authenticity of the target object can be identified, for example, preventing the use of player models as assets. On the other hand, based on the video provided by the terminal, the shooting trajectory and shooting direction are determined, thereby providing shooting suggestions. Furthermore, AR interaction can guide and control the user to shoot the scene at an accurate angle, improving the shooting quality of the target video and thus facilitating accurate target object recognition based on the video.

[0110] It should be noted that the above figures are merely illustrative of the processes included in the methods according to exemplary embodiments of this specification, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may, for example, be executed synchronously or asynchronously in multiple modules.

[0111] The following are embodiments of the apparatus described in this specification, which can be used to execute the embodiments of the methods described in this specification. For details not disclosed in the apparatus embodiments of this specification, please refer to the embodiments of the methods described in this specification.

[0112] in, Figure 11 This is a schematic diagram of a device for identifying target objects in a video according to an embodiment of this specification, which can be specifically configured in the aforementioned server 120. The device 1100 for identifying target objects in a video according to this embodiment includes: a feature point determination module 1110, a detection region determination module 1120, a repetition count determination module 1130, and an identification module 1140.

[0113] The feature point determination module 1110 is used to acquire an image sequence containing L images from the target video, wherein each image in the image sequence corresponds to a feature point reflecting three-dimensional spatial information, and L is an integer greater than 1; the detection region determination module 1120 is used to determine the detection region corresponding to the target object in the s-th image and the detection region corresponding to the target object in the (s+1)-th image, wherein each detection region corresponds to one target object, and s takes the value of an integer from 1 to L-1; the repetition quantity determination module 1130 is used to determine the repetition quantity X of the target object between the s-th image and the (s+1)-th image based on the feature points corresponding to the detection regions in the s-th image and the detection regions in the (s+1)-th image. s,s+1 ; and the aforementioned identification module 1140 is used to identify the number of repetitions X based on the aforementioned number of repetitions X s,s+1 Identify the number of the aforementioned target objects in the target video.

[0114] In an exemplary embodiment, Figure 12 A schematic diagram illustrating the structure of a device for recognizing target objects in a video, provided in accordance with another exemplary embodiment of this specification, is shown. Please refer to... Figure 11 :

[0115] In an exemplary embodiment, based on the aforementioned scheme, the s-th image contains M detection regions, and the (s+1)-th image contains N detection regions, where M and N are positive integers;

[0116] The aforementioned module 1130 for determining the number of repetitions includes: a center point determination unit 11301 and a distance calculation unit 11302.

[0117] The center point determination unit 11301 is used to: determine the center point C based on the feature points corresponding to the i-th detection region in the s-th image. i The value of i takes the value of any positive integer from 1 to M; the aforementioned center point determination unit 11301 is further configured to: determine the center point C' based on the feature points corresponding to the j-th detection region in the (s+1)-th image. j j takes the value of any positive integer from 1 to N; and the aforementioned distance calculation unit 11302 is used to: calculate the aforementioned center point C i With respect to the above center point C' j Distance D between i,j And based on the above distance D i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 .

[0118] In an exemplary embodiment, based on the foregoing scheme, the above-mentioned repetition quantity determination module 1130 further includes: a feature point clustering unit 11303.

[0119] The feature point clustering unit 11303 is configured to: cluster the feature points corresponding to the detection region in the s-th image to obtain M cluster centers corresponding to the s-th image, where M is a positive integer; the feature point clustering unit 11303 is also configured to: cluster the feature points corresponding to the detection region in the (s+1)-th image to obtain N cluster centers corresponding to the (s+1)-th image, where N is a positive integer; and the distance calculation unit 11302 is also configured to: calculate the i-th cluster center C in the s-th image. i The j-th cluster center C' in the (s+1)-th image above j Distance D between i,j Let i take the values ​​of all positive integers from 1 to M, and j take the values ​​of all positive integers from 1 to N; and according to the distance D mentioned above i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 .

[0120] In an exemplary embodiment, based on the aforementioned scheme, the s-th image contains M detection regions, and the (s+1)-th image contains N detection regions, where M and N are positive integers;

[0121] The aforementioned module 1130 for determining the number of repetitions also includes a point cloud matching unit 11304.

[0122] The point cloud matching unit 11304 is used to: obtain the sparse point cloud D corresponding to the i-th detection region in the s-th image. i And determine the above sparse point cloud D i Corresponding center point C i The value of i takes all positive integers from 1 to M; the point cloud matching unit 11304 is also used to: obtain the sparse point cloud D' corresponding to the j-th detection region in the (s+1)-th image. j And determine the above sparse point cloud D' j The corresponding center point C' j j takes the value of any positive integer from 1 to N; and the aforementioned distance calculation unit 11302 is also used to: calculate the aforementioned center point C i With respect to the above center point C' j Distance D between i,j And based on the above distance D i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 .

[0123] In an exemplary embodiment, based on the foregoing scheme, the point cloud matching unit 11304 is specifically configured to: obtain the i-th detection region contained in the first s images of the image sequence, obtain s detection regions, and match the feature points corresponding to the s detection regions to obtain the sparse point cloud D corresponding to the i-th detection region in the s-th image. i ;

[0124] The point cloud matching unit 11304 described above is further configured to: obtain the j-th detection region contained in the first s+1 images of the image sequence, obtain s+1 detection regions, and match the feature points corresponding to the s+1 detection regions to obtain the sparse point cloud D' corresponding to the j-th detection region in the s+1-th image. j .

[0125] In an exemplary embodiment, based on the foregoing scheme, the identification module 1140 includes: a authenticity identification unit 11401.

[0126] The authenticity identification unit 11401 is used for: the point cloud matching unit acquiring the sparse point cloud D corresponding to the i-th detection region in the s-th image. i Then, obtain the above sparse point cloud D. i The maximum envelope size; and the comparison between the maximum envelope size and the target envelope size, and the determination of the authenticity of the target object based on the comparison result, wherein the target envelope size is determined according to the category of the target object.

[0127] In an exemplary embodiment, based on the foregoing scheme, the aforementioned distance calculation unit 11302 is specifically used to: calculate the aforementioned distance D i,j The number of times exceeding the first preset value is determined as the number of repetitions of the target object X. s,s+1 ;

[0128] The aforementioned identification module 1140 further includes: a quantity identification unit 11402;

[0129] The quantity recognition unit 11402 is used to: calculate the sum of M and N, and then calculate the sum and X. s,s+1 The difference; and, the difference is determined as the actual number of the target objects in the s-th image and the (s+1)-th image, to obtain the number of the target objects in the target video.

[0130] In an exemplary embodiment, based on the foregoing embodiments, the target video is obtained by the terminal using motion structure recovery technology or simultaneous positioning and composition framing, and by capturing the target object using multiple sensors and cameras.

[0131] In an exemplary embodiment, based on the foregoing embodiments, the above-mentioned device 1100 further includes an image capture module 1150.

[0132] The image capture module 1150 is configured to: capture images from the target video based on one or more of the following factors to obtain the image sequence containing L images, the factors being the moving speed of the terminal, the bandwidth corresponding to the terminal, and the type of the target object.

[0133] In an exemplary embodiment, based on the foregoing embodiments, the above-mentioned device further includes: a trajectory determination module 1160, a suggestion determination module 1170, and a suggestion sending module 1180.

[0134] The trajectory determination module 1160 is used to: determine the shooting trajectory of the target object and the shooting direction of the camera based on the target video; the suggestion determination module 1170 is used to: determine the tangent direction of the target shooting point based on the shooting trajectory, calculate the angle between the tangent direction of the target shooting point and the shooting direction of the camera, and determine a shooting suggestion for the target shooting point based on the angle; and the suggestion sending module 1180 is used to: send the shooting suggestion to the terminal so as to display the shooting suggestion on the terminal.

[0135] It should be noted that the target object recognition device configured on the server provided in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0136] Figure 13 This is a schematic diagram of a device for identifying target objects in a video according to an embodiment of this specification, which can be specifically configured in the aforementioned terminal 110. The aforementioned device 1300 for identifying target objects in a video according to this embodiment includes: a video capturing module 1210 and a video sending module 1320.

[0137] The video capturing module 1310 is used to capture the target video; and the video sending module 1320 is used to send the target video to the server, so that the server:

[0138] From the target video, obtain an image sequence containing L images, where each image in the image sequence corresponds to feature points reflecting three-dimensional spatial information, and L is an integer greater than 1; determine the detection region corresponding to the target object in the s-th image and the detection region corresponding to the target object in the (s+1)-th image, where each detection region corresponds to one target object, and s takes the value of an integer from 1 to L-1; based on the feature points corresponding to the detection regions in the s-th and (s+1)-th images, determine the number of repetitions X of the target object between the s-th and (s+1)-th images. s,s+1 ; and, based on the above repetition count X s,s+1 Identify the number of the aforementioned target objects in the target video.

[0139] In an exemplary embodiment, Figure 14 A schematic diagram illustrating the structure of a device for recognizing target objects in a video, provided in accordance with another exemplary embodiment of this specification, is shown. Please refer to... Figure 13 :

[0140] In an exemplary embodiment, based on the aforementioned scheme, the video shooting module 1310 is specifically used to: capture the target video based on motion structure recovery technology or simultaneous positioning and composition frame, and based on multiple sensors and cameras to capture the target object.

[0141] In an exemplary embodiment, based on the foregoing scheme, the above-mentioned device 1300 further includes: a suggestion receiving module 1330.

[0142] The aforementioned suggestion receiving module 1330 is used to receive the shooting suggestion sent by the server after the aforementioned target video is sent to the server, and to display the shooting suggestion.

[0143] The aforementioned shooting suggestion is determined by the server based on the shooting trajectory, which determines the tangent direction of the target shooting point, and calculates the angle between the tangent direction of the target shooting point and the shooting direction of the camera. The aforementioned shooting trajectory and the shooting direction of the camera are obtained based on the aforementioned target video.

[0144] It should be noted that the target object recognition device configured on the terminal in the above embodiments is only an example of the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0145] Furthermore, the device for identifying target objects in videos and the method for identifying target objects in videos provided in the above embodiments belong to the same concept. Therefore, for details not disclosed in the device embodiments of this specification, please refer to the embodiments of the method for identifying target objects in videos described above, which will not be repeated here.

[0146] The following are system embodiments of this specification, which can be used to execute the method embodiments of this specification. For details not disclosed in the system embodiments of this specification, please refer to the method embodiments of this specification.

[0147] in, Figure 15 This is a schematic diagram of a system for recognizing target objects in a video, provided as an embodiment of this specification. (Refer to...) Figure 15 The system includes: server 120 and terminal 110.

[0148] The above-described method embodiments can be implemented through information interaction between server 120 and terminal 110.

[0149] For example, refer to Figure 16 In the asset inventory scenario, server 120 includes the following modules: asset type, type recognition, 3D object recognition, video processing, keyframe image (image in image sequence) processing, point cloud / 3D mesh visualization, point cloud densification, feature point / point cloud processing, manual approval, asset inventory, object scale restoration, interactive guidance control, and trajectory / azimuth, etc. Server 110 includes the following modules: AR interaction, front-end image / video / IMU acquisition, SLAM SDK trajectory / orientation / feature points, etc.

[0150] For example, based on the modules in server 120 and terminal 110 described above, the method embodiments described above can be implemented.

[0151] The example numbers in this specification are for descriptive purposes only and do not represent the superiority or inferiority of the examples.

[0152] This specification also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the methods described above.

[0153] Figure 17 This is a schematic diagram of the electronic device provided in the embodiments of this specification. Please refer to... Figure 10 As shown, the electronic device 1700 includes a processor 1701 and a memory 1702.

[0154] In this embodiment, processor 1701 is the control center of the computer system and can be a processor of a physical machine or a processor of a virtual machine. Processor 1701 may include one or more processing cores, such as a 4-core processor or an 8-core processor. Processor 1701 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). Processor 1001 may also include a main processor and a coprocessor; the main processor is used to process data in the wake-up state, and the coprocessor is a low-power processor used to process data in the standby state.

[0155] In the embodiments described in this specification, when the processor 1701 is configured in the server 120, it can be specifically used for:

[0156] Obtain an image sequence containing L images from the target video, where each image in the sequence corresponds to feature points reflecting three-dimensional spatial information, and L is an integer greater than 1; determine the detection region corresponding to the target object in the s-th image and the detection region corresponding to the target object in the (s+1)-th image, where each detection region corresponds to one target object, and s takes the value of an integer from 1 to L-1; based on the feature points corresponding to the detection regions in the s-th and (s+1)-th images, determine the number of repetitions X of the target object between the s-th and (s+1)-th images. s,s+1 ; and, based on the above repetition count X s,s+1 Identify the number of the aforementioned target objects in the target video.

[0157] Furthermore, the s-th image contains M detection regions, and the (s+1)-th image contains N detection regions, where M and N are positive integers;

[0158] Based on the feature points corresponding to the detection region in the s-th image and the feature points corresponding to the detection region in the (s+1)-th image, the number of repetitions X of the target object between the s-th image and the (s+1)-th image is determined. s,s+1 This includes: determining the center point C based on the feature points corresponding to the i-th detection region in the s-th image. i i takes the value of any positive integer from 1 to M; the center point C' is determined based on the feature points corresponding to the j-th detection region in the (s+1)-th image. jj takes the value of any positive integer from 1 to N; and, calculate the above center point C. i With respect to the above center point C' j Distance D between i,j And based on the above distance D i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 .

[0159] Furthermore, based on the feature points corresponding to the detection region in the s-th image and the feature points corresponding to the detection region in the (s+1)-th image, the number of repetitions X of the target object between the s-th image and the (s+1)-th image is determined. s,s+1 This includes: clustering the feature points corresponding to the detection regions in the s-th image to obtain M cluster centers for the s-th image, where M is a positive integer; clustering the feature points corresponding to the detection regions in the (s+1)-th image to obtain N cluster centers for the (s+1)-th image, where N is a positive integer; and calculating the i-th cluster center C in the s-th image. i The j-th cluster center C' in the (s+1)-th image above j Distance D between i,j Let i take the values ​​of all positive integers from 1 to M, and j take the values ​​of all positive integers from 1 to N; and according to the distance D mentioned above i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 .

[0160] Furthermore, the s-th image contains M detection regions, and the (s+1)-th image contains N detection regions, where M and N are positive integers;

[0161] Based on the feature points corresponding to the detection region in the s-th image and the feature points corresponding to the detection region in the (s+1)-th image, the number of repetitions X of the target object between the s-th image and the (s+1)-th image is determined. s,s+1 This includes: obtaining the sparse point cloud D corresponding to the i-th detection region in the s-th image above. i And determine the above sparse point cloud D i Corresponding center point C i i takes the value of any positive integer from 1 to M; obtain the sparse point cloud D' corresponding to the j-th detection region in the (s+1)-th image above. j And determine the above sparse point cloud D' j The corresponding center point C' j j takes the value of any positive integer from 1 to N; and, calculate the above center point C. i With respect to the above center point C' j Distance D between i,jAnd based on the above distance D i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 .

[0162] Furthermore, the sparse point cloud D corresponding to the i-th detection region in the s-th image is obtained as described above. i This includes: obtaining the i-th detection region contained in the first s images of the above image sequence, obtaining s detection regions, and matching the feature points corresponding to the above s detection regions to obtain the sparse point cloud D corresponding to the i-th detection region in the above s-th image. i ;

[0163] The above describes obtaining the sparse point cloud D' corresponding to the j-th detection region in the (s+1)-th image. j This includes: obtaining the j-th detection region contained in the first s+1 images of the above image sequence, obtaining s+1 detection regions, and matching the feature points corresponding to the above s+1 detection regions to obtain the sparse point cloud D' corresponding to the j-th detection region in the above s+1 image. j .

[0164] Furthermore, the aforementioned processor 1701 is specifically used for:

[0165] The above describes obtaining the sparse point cloud D corresponding to the i-th detection region in the s-th image. i Then, obtain the above sparse point cloud D. i The maximum envelope size; and the comparison between the maximum envelope size and the target envelope size, and the determination of the authenticity of the target object based on the comparison result, wherein the target envelope size is determined according to the category of the target object.

[0166] Furthermore, based on the aforementioned distance D... i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 This includes: the distance D mentioned above i,j The number of times exceeding the first preset value is determined as the number of repetitions of the target object X. s,s+1 ;

[0167] The above is based on the number of repetitions X. s,s+1 Identifying the number of target objects in the target video includes: calculating the sum of M and N, and then calculating the sum and X. s,s+1 The difference; and, the difference is determined as the actual number of the target objects in the s-th image and the (s+1)-th image, to obtain the number of the target objects in the target video.

[0168] Furthermore, the aforementioned target video is obtained by the terminal based on motion structure recovery technology or simultaneous positioning and mapping framing, as well as by capturing the aforementioned target object using multiple sensors and cameras.

[0169] Furthermore, the above-mentioned method of obtaining an image sequence containing L images from a target video includes: extracting images from the target video to obtain the image sequence containing L images based on one or more of the following factors, including the moving speed of the terminal, the bandwidth corresponding to the terminal, and the type of the target object.

[0170] Furthermore, the aforementioned processor 1701 is specifically used for:

[0171] Based on the target video, determine the shooting trajectory of the target object and the shooting direction of the camera; based on the shooting trajectory, determine the tangent direction of the target shooting point, and calculate the angle between the tangent direction of the target shooting point and the shooting direction of the camera, and based on the angle, determine a shooting suggestion for the target shooting point; and send the shooting suggestion to the terminal to display the shooting suggestion on the terminal.

[0172] In the embodiments described in this specification, when the processor 1701 is configured in the terminal 110, it can be specifically used for:

[0173] The process involves capturing a target video and sending the target video to a server, so that the server: acquires an image sequence containing L images from the target video, wherein each image in the image sequence corresponds to a feature point reflecting three-dimensional spatial information, and L is an integer greater than 1; determines the detection region corresponding to the target object in the s-th image and the detection region corresponding to the target object in the (s+1)-th image, each detection region corresponding to one target object, and s taking the value of an integer from 1 to L-1; and determines the number of repetitions X of the target object between the s-th image and the (s+1)-th image based on the feature points corresponding to the detection regions in the s-th image and the detection regions in the (s+1)-th image. s,s+1 ; and, based on the above repetition count X s,s+1 Identify the number of the aforementioned target objects in the target video.

[0174] Furthermore, the aforementioned target video capture includes: capturing the target video based on motion structure recovery technology or simultaneous positioning and composition framing, as well as capturing the target object based on multiple sensors and cameras.

[0175] Furthermore, the processor 1701 is specifically used to: after sending the target video to the server, receive the shooting suggestion sent by the server, and display the shooting suggestion;

[0176] The aforementioned shooting suggestion is determined by the server based on the shooting trajectory, which determines the tangent direction of the target shooting point, and calculates the angle between the tangent direction of the target shooting point and the shooting direction of the camera. The aforementioned shooting trajectory and the shooting direction of the camera are obtained based on the aforementioned target video.

[0177] Memory 1702 may include one or more computer-readable storage media, which may be non-transitory. Memory 1702 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments of this specification, the non-transitory computer-readable storage media in memory 1702 is used to store at least one instruction for execution by processor 1701 to implement the methods in the embodiments of this specification.

[0178] In some embodiments, the electronic device 1700 further includes a peripheral device interface 1703 and at least one peripheral device. The processor 1701, memory 1702, and peripheral device interface 1703 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1703 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of a display screen 1704, a camera 1705, and an audio circuit 1706.

[0179] Peripheral interface 1703 can be used to connect at least one input / output (I / O) related peripheral device to processor 1701 and memory 1702. In some embodiments of this specification, processor 1701, memory 1702, and peripheral interface 1703 are integrated on the same chip or circuit board; in some other embodiments of this specification, any one or two of processor 1701, memory 1702, and peripheral interface 1703 can be implemented on separate chips or circuit boards. This specification does not specifically limit the embodiments in this regard.

[0180] Display screen 1704 is used to display a user interface (UI). The UI may include graphics, text, icons, video, and any combination thereof. When display screen 1704 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1701 for processing. In this case, display screen 1704 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments of this specification, there may be one display screen 1704, which serves as the front panel of electronic device 1700; in other embodiments, there may be at least two display screens 1704, respectively disposed on different surfaces of electronic device 1700 or in a folded design; in still other embodiments, display screen 1704 may be a flexible display screen, disposed on a curved or folded surface of electronic device 1700. Furthermore, display screen 1704 may also be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. The display screen 1704 can be made of materials such as liquid crystal display (LCD) and organic light-emitting diode (OLED).

[0181] Camera 1705 is used to capture images or videos. Optionally, camera 1705 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the electronic device, and the rear-facing camera is located on the back of the electronic device. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, virtual reality (VR) shooting, or other fusion shooting functions. In some embodiments of this specification, camera 1705 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cool light flash, which can be used for light compensation at different color temperatures.

[0182] The audio circuit 1706 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input to the processor 1701 for processing. For stereo sound acquisition or noise reduction purposes, there may be multiple microphones, each located in a different part of the electronic device 1700. The microphone may also be an array microphone or an omnidirectional microphone.

[0183] Power supply 1707 is used to supply power to various components in electronic device 1700. Power supply 1707 can be alternating current, direct current, a disposable battery, or a rechargeable battery. When power supply 1707 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged via a wired line, and a wireless rechargeable battery is a battery that is charged via a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0184] The block diagrams of the electronic device shown in the embodiments of this specification do not constitute a limitation on the electronic device 1700. The electronic device 1700 may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0185] This specification also provides a computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform one or more steps in the above embodiments. If the constituent modules of the above-described table recognition and reconstruction device are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium.

[0186] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this specification are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in or transmitted through a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, Digital Subscriber Line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The aforementioned available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., Digital Versatile Discs (DVDs)), or semiconductor media (e.g., Solid State Disks (SSDs)).

[0187] In the description of this specification, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of these terms in this specification based on the specific circumstances. Furthermore, in the description of this specification, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0188] It should be noted that the above description describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims may be performed in a different order than that shown in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0189] The above description is merely a specific embodiment of this specification, but the scope of protection of this specification is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this specification should be included within the scope of protection of this specification. Therefore, equivalent variations made in accordance with the claims of this specification are still within the scope of this specification.

Claims

1. A method for identifying target objects in a video, wherein, Applied to a server, the method includes: Obtain an image sequence containing L images from the target video, wherein each image in the image sequence corresponds to a feature point reflecting three-dimensional spatial information, and L is an integer greater than 1; Determine the detection region corresponding to the target object in the s-th image of the image sequence and determine the detection region corresponding to the target object in the (s+1)-th image. Each detection region corresponds to one target object, and s takes the value of an integer from 1 to L-1. Based on the feature points corresponding to the detected region in the s-th image and the feature points corresponding to the detected region in the (s+1)-th image, determine the number X of repetitions of the target object between the s-th image and the (s+1)-th image. s,s+1 ; According to the number of repetitions X s,s+1 Identify the number of target objects in the target video; Specifically, the step of determining the number of repetitions X of the target object between the s-th image and the (s+1)-th image based on the feature points corresponding to the detection region in the s-th image and the feature points corresponding to the detection region in the (s+1)-th image. s,s+1 ,include: Cluster the feature points corresponding to the detection region in the s-th image to obtain M cluster centers corresponding to the s-th image, where M is a positive integer; Cluster the feature points corresponding to the detection region in the (s+1)th image to obtain N cluster centers corresponding to the (s+1)th image, where N is a positive integer; Calculate the i-th cluster center C in the s-th image. i The j-th cluster center C' in the (s+1)-th image j Distance D between i,j i takes the value of any positive integer from 1 to M, and j takes the value of any positive integer from 1 to N; According to the distance D i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 ; The image sequence has a chronological order; the first preset value is associated with the type of the target object.

2. The method according to claim 1, wherein, The s-th image contains M detection regions, and the (s+1)-th image contains N detection regions, where M and N are positive integers; The number of repetitions X of the target object between the s-th image and the (s+1)-th image is determined based on the feature points corresponding to the detected regions in the s-th image and the (s+1)-th image. s,s+1 ,include: The center point C is determined based on the feature points corresponding to the i-th detection region in the s-th image. i i takes the value of any positive integer from 1 to M; The center point C' is determined based on the feature points corresponding to the j-th detection region in the (s+1)-th image. j j takes the value of any positive integer from 1 to N; Calculate the center point C i With the center point C' j Distance D between i,j And according to the distance D i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 .

3. The method according to claim 1, wherein, The s-th image contains M detection regions, and the (s+1)-th image contains N detection regions, where M and N are positive integers; The number of repetitions X of the target object between the s-th image and the (s+1)-th image is determined based on the feature points corresponding to the detected regions in the s-th image and the (s+1)-th image. s,s+1 ,include: Obtain the sparse point cloud D corresponding to the i-th detection region in the s-th image. i And determine the sparse point cloud D i Corresponding center point C i i takes the value of any positive integer from 1 to M; Obtain the sparse point cloud D' corresponding to the j-th detection region in the (s+1)-th image. j And determine the sparse point cloud D' j The corresponding center point C' j j takes the value of any positive integer from 1 to N; Calculate the center point C i With the center point C' j Distance D between i,j And according to the distance D i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 .

4. The method according to claim 3, wherein, The sparse point cloud D corresponding to the i-th detection region in the s-th image is obtained. i ,include: Obtain the i-th detection region contained in the first s images of the image sequence, resulting in s detection regions. Then, match the feature points corresponding to the s detection regions to obtain the sparse point cloud D corresponding to the i-th detection region in the s-th image. i ; The sparse point cloud D' corresponding to the j-th detection region in the (s+1)-th image is obtained. j ,include: Obtain the j-th detection region contained in the first s+1 images of the image sequence to obtain s+1 detection regions, and match the feature points corresponding to the s+1 detection regions to obtain the sparse point cloud D' corresponding to the j-th detection region in the s+1-th image. j .

5. The method according to claim 3, wherein, The sparse point cloud D corresponding to the i-th detection region in the s-th image is obtained. i Subsequently, the method further includes: Obtain the sparse point cloud D i Maximum envelope size; The maximum envelope size is compared with the target envelope size, and the identification result regarding the authenticity of the target object is determined based on the comparison result. The target envelope size is determined according to the category of the target object.

6. The method according to any one of claims 1 to 5, wherein, According to the distance D i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 ,include: The distance D i,j The number of times the value exceeds the first preset value is determined as the number of repetitions of the target object X. s,s+1 ; The number of repetitions X s,s+1 Identifying the number of target objects in the target video includes: Calculate the sum of M and N, then calculate the sum and X. s,s+1 The difference; The difference is determined as the actual number of target objects in the s-th image and the (s+1)-th image, thus obtaining the number of target objects in the target video.

7. The method according to claim 1, wherein, The target video is obtained by the terminal based on motion structure recovery technology or simultaneous positioning and composition framing, and by capturing the target object using multiple sensors and cameras.

8. The method according to claim 7, wherein, The step of obtaining an image sequence containing L images from a target video includes: The image sequence containing L images is obtained by extracting images from the target video based on one or more of the following factors: the moving speed of the terminal, the bandwidth corresponding to the terminal, and the type of the target object.

9. The method according to claim 7 or 8, wherein, The method further includes: The shooting trajectory of the target object and the shooting direction of the camera are determined based on the target video; The tangent direction of the target shooting point is determined based on the shooting trajectory, and the angle between the tangent direction of the target shooting point and the shooting direction of the camera is calculated. Based on the angle, a shooting suggestion for the target shooting point is determined. The shooting suggestion is sent to the terminal so that it is displayed on the terminal.

10. A method for identifying target objects in a video, wherein, Applied to a terminal, the method includes: Shoot the target video; Send the target video to the server so that the server: An image sequence containing L images is obtained from the target video, wherein each image in the image sequence corresponds to feature points reflecting three-dimensional spatial information, and L is an integer greater than 1; the detection region corresponding to the target object in the s-th image and the detection region corresponding to the target object in the (s+1)-th image are determined, each detection region corresponding to one target object, and s takes the value of an integer from 1 to L-1; based on the feature points corresponding to the detection regions in the s-th image and the (s+1)-th image, the number of repetitions X of the target object between the s-th image and the (s+1)-th image is determined. s,s+1 ; and, according to the number of repetitions X s,s+1 Identify the number of target objects in the target video; Specifically, the step of determining the number of repetitions X of the target object between the s-th image and the (s+1)-th image based on the feature points corresponding to the detection region in the s-th image and the feature points corresponding to the detection region in the (s+1)-th image. s,s+1 ,include: Cluster the feature points corresponding to the detection region in the s-th image to obtain M cluster centers corresponding to the s-th image, where M is a positive integer; Cluster the feature points corresponding to the detection region in the (s+1)th image to obtain N cluster centers corresponding to the (s+1)th image, where N is a positive integer; Calculate the i-th cluster center C in the s-th image. i The j-th cluster center C' in the (s+1)-th image j Distance D between i,j i takes the value of any positive integer from 1 to M, and j takes the value of any positive integer from 1 to N; According to the distance D i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 ; The image sequence has a chronological order; the first preset value is associated with the type of the target object.

11. The method according to claim 10, wherein, The target video being captured includes: Based on motion structure recovery technology or simultaneous positioning and mapping framework, and based on multiple sensors and cameras to capture the target object, the target video is captured.

12. The method according to claim 10, wherein, After sending the target video to the server, the method further includes: Receive and display the shooting suggestion sent by the server; The shooting suggestion is determined by the server based on the shooting trajectory, which determines the tangent direction of the target shooting point, and calculates the angle between the tangent direction of the target shooting point and the shooting direction of the camera. The shooting trajectory and the shooting direction of the camera are obtained based on the target video.

13. A device for identifying target objects in a video, wherein, Configured on a server, the device includes: The feature point determination module is used to obtain an image sequence containing L images from the target video, wherein each image in the image sequence corresponds to a feature point reflecting three-dimensional spatial information, and L is an integer greater than 1; The detection region determination module is used to determine the detection region corresponding to the target object in the s-th image of the image sequence and to determine the detection region corresponding to the target object in the (s+1)-th image. Each detection region corresponds to one target object, and s takes the value of an integer from 1 to L-1. The repetition count determination module is used to determine the repetition count X of the target object between the s-th image and the s+1-th image based on the feature points corresponding to the detection region in the s-th image and the feature points corresponding to the detection region in the (s+1)-th image. s,s+1 ; The identification module is used to identify the number of repetitions X. s,s+1 Identify the number of target objects in the target video; Specifically, the repetition count determination module is used to cluster the feature points corresponding to the detection region in the s-th image to obtain M cluster centers corresponding to the s-th image, where M is a positive integer; to cluster the feature points corresponding to the detection region in the (s+1)-th image to obtain N cluster centers corresponding to the (s+1)-th image, where N is a positive integer; and to calculate the i-th cluster center C in the s-th image. i The j-th cluster center C' in the (s+1)-th image j Distance D between i,j i takes the value of any positive integer from 1 to M, and j takes the value of any positive integer from 1 to N; according to the distance D i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 ; The image sequence has a chronological order; the first preset value is associated with the type of the target object.

14. A device for identifying target objects in a video, wherein, Configured in a terminal, the device includes: The video recording module is used to capture video of the target. The sending module is used to send the target video to the server, so that the server: An image sequence containing L images is obtained from the target video, wherein each image in the image sequence corresponds to feature points reflecting three-dimensional spatial information, and L is an integer greater than 1; the detection region corresponding to the target object in the s-th image and the detection region corresponding to the target object in the (s+1)-th image are determined, each detection region corresponding to one target object, and s takes the value of an integer from 1 to L-1; based on the feature points corresponding to the detection regions in the s-th image and the (s+1)-th image, the number of repetitions X of the target object between the s-th image and the (s+1)-th image is determined. s,s+1 ; and, according to the number of repetitions X s,s+1 Identify the number of target objects in the target video; Specifically, the step of determining the number of repetitions X of the target object between the s-th image and the (s+1)-th image based on the feature points corresponding to the detection region in the s-th image and the feature points corresponding to the detection region in the (s+1)-th image. s,s+1 ,include: Cluster the feature points corresponding to the detection region in the s-th image to obtain M cluster centers corresponding to the s-th image, where M is a positive integer; Cluster the feature points corresponding to the detection region in the (s+1)th image to obtain N cluster centers corresponding to the (s+1)th image, where N is a positive integer; Calculate the i-th cluster center C in the s-th image. i The j-th cluster center C' in the (s+1)-th image j Distance D between i,j i takes the value of any positive integer from 1 to M, and j takes the value of any positive integer from 1 to N; According to the distance D i,j The number of repetitions X of the target object is determined by the first preset value. s,s+1 ; The image sequence has a chronological order; the first preset value is associated with the type of the target object.

15. A system for recognizing target objects in a video, wherein, The system includes: A server and a terminal, wherein the server is configured to perform the method for identifying a target object in a video as described in any one of claims 1 to 9; and the terminal is configured to perform the method for identifying a target object in a video as described in any one of claims 10 to 12.

16. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the method for identifying a target object in a video as described in any one of claims 1 to 9; or, it implements the method for identifying a target object in a video as described in any one of claims 10 to 12.

17. A computer-readable storage medium storing instructions that, when executed on a computer or processor, cause the computer or processor to perform the method for identifying a target object in a video as described in any one of claims 1 to 9; or to implement the method for identifying a target object in a video as described in any one of claims 10 to 12.