Method, device, storage medium and system for distinguishing target object

By acquiring the robot's movement trajectory information within the target area and combining visual fusion and spatiotemporal modal information for multimodal feature fusion, the problem of low accuracy in robot differentiation results in existing technologies is solved, achieving higher differentiation accuracy.

CN114595745BActive Publication Date: 2026-01-02ALIBABA (CHINA) CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202210102025.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-27
Publication Date
2026-01-02
Estimated Expiration
2042-01-27

AI Technical Summary

Technical Problem

In existing technologies, when distinguishing robots by directly matching their features with video acquisition devices, the accuracy of the distinction results is low and erroneous matching is prone to occur.

Method used

By acquiring the movement trajectory information of the target object within the target area, and combining visual fusion information and spatiotemporal modal information to perform multimodal feature fusion, a more accurate differentiation result can be obtained.

Benefits of technology

This improves the accuracy of the robot's differentiation results and solves the problem of low differentiation accuracy in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114595745B_ABST
    Figure CN114595745B_ABST
Patent Text Reader

Abstract

The application discloses a method, device, storage medium and system for distinguishing target objects. The method comprises the following steps: acquiring movement track information of at least one target object in a target area, wherein the target area is used for determining the activity range of the at least one target object; acquiring visual fusion information and space-time mode information of the at least one target object based on the movement track information; and obtaining a distinguishing result by using the visual fusion information and the space-time mode information, wherein the distinguishing result is used for distinguishing each target object in the at least one target object. The application solves the technical problem of low accuracy of the distinguishing result of the processing method for directly distinguishing objects based on the collected features in the related art.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a method, device, storage medium and system for distinguishing target objects. BACKGROUND

[0002] In business scenarios such as logistics, e-commerce, catering, and medical treatment, more and more robots are put into business operation. In order to control the business activities, multiple video acquisition devices are usually included in these business scenarios. The video acquisition devices can be used to acquire the activity trajectories of multiple robots in the scene. However, when the business activities are further analyzed and regulated, it is necessary to distinguish the multiple robots in the scene to complete the data analysis in the entire scene.

[0003] In related solutions, the method for distinguishing robots by visual recognition mainly includes: capturing a robot by a video acquisition device and extracting features of the robot; directly matching the features of the robot with existing feature data in a database; and distinguishing different robots based on the matching result. However, this method has the following defects: the feature modality used for matching is single, similar robots are easily matched by mistake, and the accuracy of the distinguishing result is low.

[0004] At present, there is no effective solution to the above problems. SUMMARY

[0005] Embodiments of the present application provide a method, device, storage medium and system for distinguishing target objects, to at least solve the technical problem of low accuracy of the distinguishing result in the processing method for distinguishing objects directly based on the collected features in related technologies.

[0006] According to an aspect of an embodiment of the present application, a method for distinguishing target objects is provided, including: acquiring movement trajectory information of at least one target object in a target area, wherein the target area is used to determine the activity range of the at least one target object; based on the movement trajectory information, acquiring visual fusion information and spatio-temporal modality information of the at least one target object; and using the visual fusion information and the spatio-temporal modality information to obtain a distinguishing result, wherein the distinguishing result is used to distinguish each target object in the at least one target object.

[0007] According to another aspect of an embodiment of the present application, a method for distinguishing target objects is also provided, including: receiving movement trajectory information of at least one target object in a target area from a client, wherein the target area is used to determine the activity range of the at least one target object; based on the movement trajectory information, acquiring visual fusion information and spatio-temporal modality information of the at least one target object, and using the visual fusion information and the spatio-temporal modality information to obtain a distinguishing result, wherein the distinguishing result is used to distinguish each target object in the at least one target object; and feeding back the distinguishing result to the client.

[0008] According to a further aspect of the embodiments of the present application, there is also provided an apparatus for distinguishing target objects, comprising: a first obtaining module, configured to obtain movement trajectory information of at least one target object in a target area, wherein the target area is used to determine an activity range of the at least one target object; a second obtaining module, configured to obtain visual fusion information and spatio-temporal modality information of the at least one target object based on the movement trajectory information; and a distinguishing module, configured to obtain a distinguishing result by using the visual fusion information and the spatio-temporal modality information, wherein the distinguishing result is used to distinguish each of the at least one target object.

[0009] According to a further aspect of the embodiments of the present application, there is also provided a storage medium comprising a stored program, wherein the program, when executed, controls a device in which the storage medium is located to perform any of the methods for distinguishing target objects.

[0010] According to a further aspect of the embodiments of the present application, there is also provided a processor configured to execute a program, wherein the program, when executed, performs any of the methods for distinguishing target objects.

[0011] According to a further aspect of the embodiments of the present application, there is also provided a system for distinguishing target objects, comprising: a processor; and a memory connected to the processor, configured to provide the processor with instructions for processing the following processing steps: obtaining movement trajectory information of at least one target object in a target area, wherein the target area is used to determine an activity range of the at least one target object; obtaining visual fusion information and spatio-temporal modality information of the at least one target object based on the movement trajectory information; and obtaining a distinguishing result by using the visual fusion information and the spatio-temporal modality information, wherein the distinguishing result is used to distinguish each of the at least one target object.

[0012] In the embodiments of the present application, by obtaining movement trajectory information of at least one target object in a target area, wherein the target area is used to determine an activity range of the at least one target object, and obtaining visual fusion information and spatio-temporal modality information of the at least one target object based on the movement trajectory information, and by obtaining a distinguishing result by using the visual fusion information and the spatio-temporal modality information, wherein the distinguishing result is used to distinguish each of the at least one target object, the purpose of obtaining an object distinguishing result by using multi-modality feature fusion information is achieved, thereby achieving the technical effect of improving the accuracy of the object distinguishing result, and further solving the technical problem of low accuracy of the object distinguishing result in the processing method of directly distinguishing objects based on collected features in the related art. BRIEF DESCRIPTION OF DRAWINGS

[0013] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and together with the description serve to explain the application. In the drawings:

[0014] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a method of distinguishing target objects is shown;

[0015] Figure 2 A flowchart of a method of distinguishing target objects according to an embodiment of the application is shown;

[0016] Figure 3 A schematic diagram of an optional regional camera distribution according to an embodiment of the application is shown;

[0017] Figure 4 A schematic diagram of the result of an optional regional division according to an embodiment of the application is shown;

[0018] Figure 5 A schematic diagram of an optional visual fusion similarity calculation process according to an embodiment of the application is shown;

[0019] Figure 6 A schematic diagram of an optional process of multi-modal feature fusion according to an embodiment of the application is shown;

[0020] Figure 7 A flowchart of a method of distinguishing target objects according to an embodiment of the application is shown;

[0021] Figure 8 A schematic diagram of a cloud server for distinguishing target objects according to an embodiment of the application is shown;

[0022] Figure 9 A structural schematic diagram of a device for distinguishing target objects according to an embodiment of the application is shown;

[0023] Figure 10 A structural block diagram of another computer terminal according to an embodiment of the application is shown. DETAILED DESCRIPTION

[0024] In order to make the personnel in the technical field better understand the application scheme, the technical solutions in the embodiments of the application will be described clearly and completely below in conjunction with the drawings in the embodiments of the application. Obviously, the described embodiments are only a part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the application.

[0025] It should be noted that the terms "first", "second", and the like in the description and claims of the application and the above drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example: a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0026] First, some nouns or terms that appear in the description of the embodiments of the application are applicable to the following explanations:

[0027] Modality: the characteristics of one dimension belong to a modality, for example: time characteristics, spatial characteristics belong to a modality respectively.

[0028] Similarity: refers to the degree of association between data. The higher the degree of association, the higher the similarity between the two data.

[0029] Space-time mask: refers to using selected time constraint relationship and space constraint relationship to partially or completely shield the image to be processed, to control the processing area or processing process of the image.

[0030] Embodiment 1

[0031] According to the embodiments of the application, a method for distinguishing target objects is also provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.

[0032] The method provided by the embodiment of the application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for distinguishing target objects is shown. As Figure 1As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, a keyboard, a cursor control device (such as a mouse), an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0033] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of the present invention, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0034] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for distinguishing target objects in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned method for distinguishing target objects. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0035] The transmission device 106 is configured to receive or send data via a network. The network can include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network interface controller (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module that is configured to communicate with the Internet wirelessly.

[0036] The display can be a liquid crystal display (LCD) that is touch screen, for example, which can enable a user to interact with a user interface of the computer terminal 10 (or mobile device).

[0037] It is noted that in some alternative embodiments, the above Figure 1 The computer device (or mobile device) can include hardware elements (including circuitry), software elements (including computer code stored on a computer readable medium), or a combination of both hardware and software elements. It should be noted that in some embodiments, the functions of the computer device (or mobile device) can be performed by software and / or firmware that is executed by the computer device (or mobile device). Figure 1 is merely one example of a particular implementation and is intended to illustrate the types of components that can be present in the computer device (or mobile device) described above.

[0038] In the above operating environment, the present application provides a method for distinguishing target objects as shown in Figure 2 Figure 2 is a flowchart of a method for distinguishing target objects according to an embodiment of the present application, as shown in Figure 2

[0039] At step S202, movement trajectory information of at least one target object in a target area is obtained, wherein the target area is used to determine an activity range of the at least one target object.

[0040] At step S204, visual fusion information and spatiotemporal modality information of the at least one target object are obtained based on the movement trajectory information.

[0041] At step S206, a distinguishing result is obtained using the visual fusion information and the spatiotemporal modality information, wherein the distinguishing result is used to distinguish each target object in the at least one target object.

[0042] ​​Optionally, the target area can be used to determine the activity range of the at least one target object, and the movement track information of the at least one target object in the target area can be acquired by a video acquisition device (such as a camera). For example, the target area can be an area where a robot is used to carry out business activities, such as a shopping mall, a supermarket, a hospital, a bank, a restaurant, a hotel, and a logistics distribution area, and the target object can be a household robot (such as a sweeping robot, a dishwashing robot, and an intelligent assistant), a service robot, a guiding robot, and a delivery robot. The movement track information can be the track information of a plurality of robots carrying out business activities, and the track information can include a plurality of tracks.

[0043] Optionally, based on the movement track information of the at least one target object in the target area, the visual fusion information and the spatiotemporal modal information of the at least one target object can be acquired. The visual fusion information can be information obtained by fusing feature information acquired by a plurality of video acquisition devices in the target area. The spatiotemporal modal information can be modal information determined according to a time constraint relationship and a spatial constraint relationship determined according to the movement of the target object in the target area in the real scene.

[0044] For example, the visual fusion information of a service robot in a restaurant can be the fusion features of the appearance features, size features, and behavior features of the robot acquired and extracted by a plurality of cameras in the restaurant. The corresponding time constraint relationship of the service robot can be that the time difference between two tracks of the same robot cannot be less than the shortest reachable time between the two tracks, and the corresponding spatial constraint relationship can be that two tracks appearing at the same time in the same area cannot belong to the same robot.

[0045] Optionally, the above-mentioned visual fusion information and the above-mentioned spatiotemporal modal information can be used to obtain the above-mentioned distinguishing result. The distinguishing result can be used to distinguish each target object in the at least one target object in the above-mentioned target area.

[0046] In the embodiment of the present application, by acquiring the movement track information of the at least one target object in the target area, the target area is used to determine the activity range of the at least one target object, and based on the movement track information, the visual fusion information and the spatiotemporal modal information of the at least one target object are acquired, and the distinguishing result is obtained by using the visual fusion information and the spatiotemporal modal information, wherein the distinguishing result is used to distinguish each target object in the at least one target object, so as to achieve the purpose of obtaining the object distinguishing result by using the multi-modal feature fusion information, thereby realizing the technical effect of improving the accuracy of the object distinguishing result, and further solving the technical problem that the object distinguishing result accuracy of the processing method based on the acquired features in the related art is low.

[0047] Optionally, the method for distinguishing target objects provided in the present application can be applied in, but not limited to, the application scenarios of regional digitization, robot tracking, robot identification, etc. in the fields of logistics, medical treatment, e-commerce, offline retail and wholesale, life service, etc. In the above application scenarios, more accurate robot distinguishing results can be obtained through the method of multi-modal feature fusion based on the robot feature data collected and extracted by the video collection device, so as to realize flexible control of business activities in the above application scenarios.

[0048] In an optional embodiment, in step S202, the movement trajectory information of the at least one target object in the target region is obtained, including the following method steps:

[0049] In step S221, the distribution information of the plurality of cameras in the target region and the visual range of each camera in the plurality of cameras are obtained.

[0050] In step S222, the target region is divided in the target scale based on the distribution information and the visual range, and a plurality of sub-regions are obtained.

[0051] In step S223, for each sub-region in the plurality of sub-regions, at least one camera selected in each sub-region is used to track the at least one target object, and the tracking identifier of each target object in the at least one target object is determined.

[0052] In step S224, the movement trajectory information of the at least one target object in the target region is obtained by using the tracking identifier of each target object in the at least one target object.

[0053] Optionally, the target region can be used to determine the activity range of the at least one target object. The target region can include a plurality of cameras. The distribution information of the plurality of cameras in the target region and the visual range of each camera in the plurality of cameras can be used to divide the target region in the target scale based on the distribution information and the visual range, and the plurality of sub-regions are obtained. The distribution information can be the position information of the plurality of cameras in the target region, and the visual range of each camera in the plurality of cameras can be the range covered by the photographable area of the camera.

[0054] Optionally, for each sub-region in the plurality of sub-regions, at least one camera selected in the sub-region can be used to track the at least one target object, and the tracking identifier of each target object in the at least one target object is determined. For example, in a certain sub-region of a restaurant, the at least one target object can be a plurality of service robots, and the tracking identifier can be a unique identifier corresponding to each service robot in the plurality of service robots in the tracking process of the plurality of cameras.

[0055] Optionally, the tracking identifier of each of the at least one target object can be used to obtain the movement trajectory information of the at least one target object in the target area. For example, in the entire restaurant area, according to the unique identifier of each robot in the plurality of service robots, the plurality of robots can be determined to have a plurality of trajectories when performing business activities.

[0056] For example, when distinguishing the plurality of guide robots active in the shopping mall area Area1, the method provided in the embodiment can be used. Figure 3 is a schematic diagram of an optional area camera distribution according to an embodiment of the application, as shown in Figure 3 The Cn cameras in the shopping mall area Area1 (equivalent to the target area) can correspond to different shooting directions. According to the installation positions of the Cn cameras, the camera distribution information Data1 of the shopping mall area Area1 can be obtained. According to the camera distribution information Data1 and the shooting directions of the Cn cameras, the viewable range of each of the Cn cameras can be determined, denoted as Data_view.

[0057] Figure 4 is a schematic diagram of an optional area division result according to an embodiment of the application, as shown in Figure 4 According to the camera distribution information Data1 of the Cn cameras in the shopping mall area Area1 and the viewable range Data_view of each of the N cameras, the area division scale (equivalent to the target scale) can be determined. According to the area division scale, the shopping mall area Area1 is divided into 12 sub-areas, denoted as B1-B12.

[0058] For each of the 12 sub-areas B1-B12, the plurality of guide robots active in the area can be tracked using the plurality of cameras in the area (for example, the sub-area B4 includes 6 cameras, and the sub-area B5 includes 5 cameras), and a tracking identifier (for example, robot0001, robot0002, etc.) can be determined for each of the plurality of guide robots.

[0059] For each of the 12 sub-areas B1-B12, the tracking identifier of each of the plurality of guide robots active in the area can be used to determine a plurality of movement trajectories of the guide robot. For example, the guide robot with the tracking identifier robot0001 passes through the sub-area B4 three times, and three trajectories of the guide robot in the sub-area B4 can be obtained. For another example, the guide robot with the tracking identifier robot0002 passes through the sub-area B4 twice, and two trajectories of the guide robot in the sub-area B4 can be obtained.

[0060] It should be noted that the movement trajectory information provided by the present application can include all or part of the plurality of trajectories of the guiding robot obtained in all or part of the sub-regions of the shopping mall region Area 1. The sub-region can be a sub-region of different scales.

[0061] In an optional embodiment, in step S204, based on the movement trajectory information, the visual fusion information of the at least one target object is obtained, including the following method steps:

[0062] In step S241, image tracking is performed on the at least one target object to obtain a set of to-be-processed images of each target object in the at least one target object.

[0063] In step S242, target appearance features and target body features of the at least one target object are obtained based on the set of to-be-processed images.

[0064] In step S243, the visual fusion information of the at least one target object is obtained by using the target appearance features and the target body features.

[0065] Optionally, the set of to-be-processed images of each target object in the at least one target object can be obtained by using a camera in the target region to perform image tracking on the at least one target object. The set of to-be-processed images can include a plurality of images captured by the camera and including the target object. For example, the target object can be a household robot (such as a sweeping robot, a dishwashing robot, an intelligent assistant, etc.), a service robot, a guiding robot, a delivery robot, etc. The movement trajectory information can be trajectory information of a plurality of robots when performing business activities in the target region. The trajectory information can be trajectory information recorded by a plurality of images collected by a plurality of cameras in the target region.

[0066] Optionally, based on the set of to-be-processed images, the target appearance features and the target body features of the at least one target object can be obtained. By using the target appearance features and the target body features, the visual fusion information of the at least one target object can be obtained. The visual fusion information can be information obtained based on a plurality of images collected by a plurality of video collection devices in the target region. For example, the target object can be a robot, for each robot, the target appearance features are appearance features obtained based on a plurality of images of the robot collected by a plurality of video collection devices in the target region, the target body features are body features obtained based on a plurality of images of the robot collected by a plurality of video collection devices in the target region, and by using the target appearance features and the target body features, the visual fusion information of the robot can be obtained.

[0067] In an optional embodiment, in step S242, based on the set of to-be-processed images, the target appearance features and the target body features of the at least one target object are obtained, including the following method steps:

[0068] In step S2421, feature extraction is performed on the to-be-processed images to obtain the appearance feature and the body feature corresponding to each trajectory included in the movement trajectory information.

[0069] In step S2422, the target appearance feature is calculated by using the appearance feature corresponding to each trajectory, and the target body feature is calculated by using the body feature corresponding to each trajectory.

[0070] Optionally, the to-be-processed images can be a plurality of images obtained by performing image tracking on a plurality of target objects in the target region by a plurality of cameras in the target region. The movement trajectory information can be trajectory information of the plurality of target objects when moving in the target region. The trajectory information can include a plurality of trajectories. Feature extraction is performed on the to-be-processed images to obtain the appearance feature and the body feature corresponding to each trajectory included in the movement trajectory information. The target appearance feature is calculated by using the appearance feature corresponding to each trajectory, and the target body feature is calculated by using the body feature corresponding to each trajectory.

[0071] For example, the target object can be a robot, and the plurality of modal features of the robot can include a body feature, a size feature, an appearance feature, a type feature, and the like. Based on each image in a to-be-processed image set of a plurality of robots collected by a plurality of cameras in a certain region, the appearance feature and the body feature of each trajectory in movement trajectory information of the plurality of robots collected by the plurality of cameras can be obtained through feature extraction. Then, the target appearance feature and the target body feature corresponding to each trajectory can be calculated, wherein the target appearance feature is a feature calculated by performing calculation on a plurality of appearance features extracted from a plurality of images in the to-be-processed image set of the robot, and the target body feature is a feature calculated by performing calculation on a plurality of body features extracted from a plurality of images in the to-be-processed image set of the robot.

[0072] In an optional embodiment, in step S243, the visual fusion information of at least one target object is obtained by using the target appearance feature and the target body feature, including the following method steps:

[0073] In step S2431, a first similarity matrix between trajectories included in the movement trajectory information is determined by using the target appearance feature, and a second similarity matrix between trajectories included in the movement trajectory information is determined by using the target body feature, wherein the first similarity matrix is an appearance similarity matrix, and the second similarity matrix is a body similarity matrix.

[0074] In step S2432, a first mask is calculated based on the first similarity matrix and the second similarity matrix, wherein the first mask is a visual fusion mask.

[0075] Step S2433, the visual fusion information is calculated by using the first similarity matrix and the first mask, wherein the visual fusion information is used to represent the visual fusion similarity of the at least one target object.

[0076] Optionally, the moving track information can be track information of the plurality of target objects moving in the target area. The track information can include a plurality of tracks. The first similarity matrix can be a appearance similarity matrix, and the second similarity matrix can be a shape similarity matrix. The first similarity matrix between each track included in the moving track information can be determined by using the target appearance feature, and the second similarity matrix between each track included in the moving track information can be determined by using the target shape feature.

[0077] Optionally, the first mask can be a visual fusion mask. The visual fusion information can be information obtained based on a plurality of images collected by a plurality of cameras in the target area, and can be used to represent the visual fusion similarity of the at least one target object. The first mask can be obtained by calculation based on the first similarity matrix and the second similarity matrix. The visual fusion information can be obtained by calculation by using the first similarity matrix and the first mask.

[0078] For example, when distinguishing a plurality of guide robots moving in a shopping mall area Area1, the method provided in the embodiment can be used. Figure 5 is a schematic diagram of an optional visual fusion similarity calculation process according to an embodiment of the present application, as Figure 5 shown, for a sub-area in the shopping mall area Area1, each track of N tracks obtained by image tracking of a plurality of guide robots moving in the sub-area by a plurality of cameras in the sub-area can include a plurality of images (equivalent to the above-mentioned image set to be processed). Based on the plurality of images in each track, the average facial feature (equivalent to the above-mentioned target appearance feature) and the average body feature (equivalent to the above-mentioned target shape feature) of the guide robot corresponding to the track can be obtained, which can include the following method steps:

[0079] First step: feature extraction based on the plurality of images included in the track to obtain a plurality of facial feature data of the guide robot corresponding to the track.

[0080] Second step: average processing of the plurality of facial feature data of the guide robot corresponding to the track to obtain the average facial feature of the guide robot.

[0081] Third step: selecting images with an aspect ratio less than 2 from the plurality of images included in the track to obtain remaining partial images, and checking the number of images in the partial images.

[0082] Fourthly, if the image quantity of the partial image is greater than 0, feature extraction is performed based on the partial image to obtain a plurality of body feature data of the guide robot corresponding to the trajectory, and average processing is performed on the plurality of body feature data to obtain an average body feature of the guide robot. If the image quantity of the partial image is equal to 0, feature extraction is performed based on a plurality of images included in the trajectory to obtain a plurality of body feature data of the guide robot corresponding to the trajectory, and average processing is performed on the plurality of body feature data to obtain an average body feature of the guide robot.

[0083] It should be noted that in the above method steps, images with an aspect ratio less than 2 are regarded as half-body images of the guide robot, and images with an aspect ratio greater than or equal to 2 are regarded as full-body images of the guide robot. When extracting the body feature data of the guide robot corresponding to the trajectory, the extraction is preferentially performed based on the full-body image of the guide robot, and for the guide robot for which no full-body frame is collected, the extraction is performed based on the collected half-body image.

[0084] Still as shown in Figure 5 After obtaining the average face feature and the average body feature of the guide robot corresponding to each trajectory in the N trajectories of a certain sub-region, the robot face similarity visual_face_sim_matrix (equivalent to the first similarity matrix described above) and the robot body similarity visual_body_sim_matrix (equivalent to the second similarity matrix described above) between any two trajectories in the N trajectories are further calculated.

[0085] It should be noted that the robot face similarity visual_face_sim_matrix is an N x N matrix, and the i-th row and j-th column in the matrix represent the robot face similarity between the i-th trajectory and the j-th trajectory in the N trajectories. Similarly, the robot body similarity visual_body_sim_matrix is also an N x N matrix, and the i-th row and j-th column in the matrix represent the robot body similarity between the i-th trajectory and the j-th trajectory in the N trajectories.

[0086] Still as shown in Figure 5 For a certain sub-region in the shopping area Area1, calculating the corresponding visual fusion mask merge_mask (equivalent to the first mask described above) can include the following method steps:

[0087] Firstly, an N x N matrix is generated as the visual fusion mask merge_mask, and the initial value of the matrix is all 0.

[0088] Second step, traverse the N*N elements in the visual_body_sim_matrix matrix, for the position whose element value is greater than 0.45, assign the value of the corresponding position in the merge_mask to 1.

[0089] Third step, traverse the N*N elements in the visual_face_sim_matrix matrix, for the position whose element value is greater than 0.7, assign the value of the corresponding position in the merge_mask to 1.

[0090] The above-mentioned merge_mask is used as a calculation result, which is used to further calculate the visual_merge_sim_matrix according to the following formula (1):

[0091] visual_merge_sim_matrix = visual_face_sim_matrix * merge_mask formula (1)

[0092] It should be noted that in the above method steps, the value of the i-th row and j-th column in the merge_mask is 0, which means that the i-th trajectory and the j-th trajectory in the above-mentioned N trajectories belong to different two guide robots; the value of the i-th row and j-th column in the merge_mask is 1, which means that the i-th trajectory and the j-th trajectory in the above-mentioned N trajectories may belong to the same guide robot.

[0093] It should be noted that in the above method steps, the first threshold (0.45) is used to ensure that "the robot body similarity corresponding to two trajectories needs to be greater than the threshold, and the two trajectories may belong to the same guide robot". The second threshold (0.7) is used to ensure that "no matter what the robot body similarity corresponding to two trajectories is, if the robot face similarity is greater than the threshold, the two trajectories are considered to belong to the same guide robot". For example: if the robot face similarity corresponding to two trajectories is greater than 0.7, even if the corresponding robot body similarity is less than 0.45, it is still considered that the two trajectories are considered to belong to the same guide robot. Therefore, the condition for judging whether two trajectories belong to the same guide robot is that the corresponding robot body similarity is greater than the first threshold, or the corresponding robot face similarity is greater than the second threshold.

[0094] It should be noted that the values of the first threshold and the second threshold can be flexibly selected according to different actual application scenarios. For example, in a scene with moderate scene light brightness (a scene in which the video acquisition device can acquire a higher quality image), the first threshold and the second threshold can take lower values. Conversely, in a scene with excessively bright or excessively dark scene light brightness (a scene in which the video acquisition device acquires a poor quality image), the first threshold and the second threshold can take higher values.

[0095] It is easy to note that, according to the method provided in the embodiment, the distinguishing of the target object based on the visual fusion mask considers the features of the two modalities of appearance and shape, and compared with the method of using only the features of a single modality to identify and distinguish objects in related solutions, the accuracy of the distinguishing result can be improved, which is beneficial to application in actual scenarios.

[0096] In an optional embodiment, in step S204, the spatiotemporal modality information of the at least one target object is obtained based on the movement trajectory information, including the following method steps:

[0097] In step S244, a second mask is generated according to the number of trajectories contained in the movement trajectory information, and an initial value of the second mask is set, wherein the second mask is a spatiotemporal mask.

[0098] In step S245, the initial value is adjusted based on the spatiotemporal constraint relationship between any two trajectories contained in the movement trajectory information, to obtain a target value of the second mask, so as to determine the spatiotemporal modality information, wherein the spatiotemporal modality information is used to represent the spatiotemporal modality features of the at least one target object.

[0099] Optionally, the second mask can be a spatiotemporal mask. The movement trajectory information can be trajectory information of the movement of the plurality of target objects in the target region. The trajectory information can include a plurality of trajectories. According to the number of trajectories contained in the movement trajectory information, the second mask can be generated, and an initial value of the second mask can be set.

[0100] Optionally, there is a spatiotemporal constraint relationship between each two trajectories in the plurality of trajectories contained in the movement trajectory information. Based on the spatiotemporal constraint relationship between any two trajectories contained in the movement trajectory information, the initial value of the second mask can be adjusted, and then the target value of the second mask can be obtained. The target value of the second mask can be used to determine the spatiotemporal modality information of the at least one target object in the target region. The spatiotemporal modality information can be used to represent the spatiotemporal modality features of the at least one target object.

[0101] In an optional embodiment, in step S206, the distinguishing result is obtained by using the visual fusion information and the spatiotemporal modality information, including the following method steps:

[0102] In step S261, the target fusion similarity is calculated using the visual fusion information and the spatio-temporal modal information, wherein the target fusion similarity is a multi-modal feature fusion similarity of at least one target object.

[0103] In step S262, the distinguishing result is determined based on the target fusion similarity.

[0104] Optionally, the target fusion similarity can be the multi-modal feature fusion similarity of the at least one target object. Based on the visual fusion information and the spatio-temporal modal information, the target fusion similarity can be calculated. The higher the target fusion similarity is, the more likely it is that the two trajectories corresponding to the target fusion similarity belong to the same target object.

[0105] Optionally, the distinguishing result can be determined based on the target fusion similarity. The distinguishing result can be used to distinguish each of the at least one target object in the target region. The numerical value of the target fusion similarity can be used to represent the possibility that the two trajectories corresponding to the target fusion similarity belong to the same target object. Therefore, a threshold value can be preset to determine whether the two trajectories belong to the same target object, and the numerical value of the target fusion similarity corresponding to any two trajectories in the moving trajectory information is compared. If the numerical value of the target fusion similarity is not less than the preset threshold value, it is considered that the two trajectories correspond to the same target object; if the numerical value of the target fusion similarity is less than the preset threshold value, it is considered that the two trajectories correspond to two different target objects.

[0106] For example, when distinguishing a plurality of guide robots moving in a shopping mall region Area 1, the method provided in the embodiment can be used. Figure 6 is a schematic diagram of an optional multi-modal feature fusion process according to an embodiment of the application, as shown in Figure 6 For a sub-region in the shopping mall region Area 1, a plurality of cameras in the sub-region perform image tracking on a plurality of guide robots moving in the sub-region to obtain N trajectories. Based on the N trajectories and the spatio-temporal constraint relationship between any two trajectories in the N trajectories, a spatio-temporal mask (equivalent to the second mask) is calculated, including the following method steps:

[0107] First, an N*N matrix is generated as a spatio-temporal mask st_mask, and the initial value of the matrix is all 1.

[0108] The second step is to iterate through the N×N elements of the spatiotemporal mask st_mask matrix. For each (i, j) position, check the timestamps of the corresponding two trajectories (the i-th and j-th trajectories among the N trajectories). If the two trajectories have appeared simultaneously in any sub-region, then the value of the (i, j) position is set to 0.

[0109] The third step is to iterate through the N×N elements of the spatiotemporal mask st_mask matrix, and for each (i, j) position, calculate the time difference Δt between the corresponding two trajectories (the i-th and j-th trajectories out of the N trajectories). ij And the shortest reachable time t between the two trajectories min Compare the time difference Δt ij With the shortest reach time t min The size relationship, if t min <Δt ij If so, the values ​​of the (i, j) position and its symmetrical position (j, i) of the spatiotemporal mask st_mask are both assigned to 0.

[0110] It should be noted that the time difference between the i-th and j-th trajectories in the N trajectories is Δt. ij Δt represents the time range between the end time t1 of the i-th and j-th trajectories and the beginning time t2 of the subsequent trajectories. ij = t2-t1.

[0111] It should be noted that the shortest reachability time t between the i-th and j-th trajectories in the N trajectories is... min The calculation method is as follows: if the sub-region where the i-th trajectory is located is adjacent to the sub-region where the j-th trajectory is located, then the shortest reachable time t is... min =0; If the sub-region containing the i-th trajectory is not adjacent to the sub-region containing the j-th trajectory, then obtain the distance x meters between the two sub-regions, and the shortest reachable time. Seconds (assuming the robot moves at a constant speed of 2 meters per second).

[0112] The spatiotemporal mask st_mask mentioned above is used as the calculation result, which is then used to further calculate the final multimodal feature fusion similarity merge_sim_matrix according to the following formula (2):

[0113] merge_sim_matrix=visual_merge_sim_matrix×st_mask formula (2)

[0114] It should be noted that in the above method steps, the value of the i-th row and the j-th column in the space-time mask st_mask is 0, which indicates that the i-th trajectory and the j-th trajectory in the N trajectories belong to different guide robots; and the value of the i-th row and the j-th column in the visual fusion mask merge_mask is 1, which indicates that the i-th trajectory and the j-th trajectory in the N trajectories may belong to the same guide robot.

[0115] It should be noted that in the above method steps, the time constraint relationship is that the time difference between two trajectories of the same robot cannot be less than the shortest reachable time between the two trajectories, and the space constraint relationship can be that two trajectories appearing at the same time in the same area cannot belong to the same robot. Through the above method of calculating the space-time mask, all two trajectories that do not satisfy the time constraint relationship and the space constraint relationship can be marked as not belonging to the same guide robot.

[0116] It is easy to note that according to the method provided by the embodiment, the target object is distinguished based on the multi-modal feature fusion similarity, the appearance feature and the shape feature in the vision are considered, and the space-time feature is also considered, which can further distinguish objects with similar appearance features and shape features. Compared with the method of using only single modal features to recognize and distinguish objects in related schemes, the method provided by the embodiment can improve the accuracy of the distinguishing result, which is beneficial to application in actual scenes.

[0117] It should be noted that in the method provided by the embodiment, the plurality of modal features for fusion can include but are not limited to the appearance feature, the shape feature and the space-time feature. For example, in different application scenarios, the plurality of modal features can also include category features, brand features, action features, etc.

[0118] An embodiment of the present application also provides a method for distinguishing target objects, which runs on a cloud server, Figure 7 is a flowchart of an optional method for distinguishing target objects according to an embodiment of the present application, as Figure 7 shown, the method for distinguishing target objects comprises:

[0119] Step S702, receiving movement trajectory information of at least one target object in a target area from a client, wherein the target area is used to determine the activity range of the at least one target object;

[0120] Step S704, obtaining visual fusion information and space-time modal information of the at least one target object based on the movement trajectory information, and obtaining a distinguishing result by using the visual fusion information and the space-time modal information, wherein the distinguishing result is used to distinguish each target object in the at least one target object;

[0121] Step S706, the result of the distinguishing is fed back to the client.

[0122] Optionally, Figure 8 is a schematic diagram of distinguishing target objects in a cloud server according to an embodiment of the present application, as shown in Figure 8 The client uploads movement trajectory information of at least one target object in a target area to the cloud server, wherein the target area is used to determine the activity range of the at least one target object; the cloud server obtains visual fusion information and space-time modal information of the at least one target object based on the movement trajectory information, and obtains a distinguishing result by using the visual fusion information and the space-time modal information, wherein the distinguishing result is used to distinguish each target object in the at least one target object. Then, the cloud server feeds back the distinguishing result to the client, and the final distinguishing result is provided to the user through the graphical user interface of the client.

[0123] It should be noted that the above-mentioned method for distinguishing target objects provided by the embodiments of the present application can be applied to, but is not limited to, the practical application scenarios of regional digitization, robot tracking, robot identification, etc. in the fields of logistics, medical treatment, e-commerce, offline retail and wholesale, life service, etc. The SaaS server and the client are interacted, the robot feature data collected and extracted by the video acquisition device is used, the more accurate robot distinguishing result is obtained by the method of multi-modal feature fusion, and the returned distinguishing result is provided to the user through the client.

[0124] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0125] From the above description of the embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software and the necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, or network device, etc.) execute the methods described in the embodiments of the present application.

[0126] Embodiment 2

[0127] According to an embodiment of the present application, a device for implementing the method of distinguishing target objects is also provided, Figure 9 is a structural schematic diagram of a device for distinguishing target objects according to an embodiment of the present application, as Figure 9 shown, the device comprises a first acquisition module 901, a second acquisition module 902, and a distinguishing module 903, wherein,

[0128] The first acquisition module 901 is configured to acquire movement trajectory information of at least one target object in a target area, wherein the target area is used to determine an activity range of the at least one target object; the second acquisition module 902 is configured to acquire visual fusion information and spatio-temporal modality information of the at least one target object based on the movement trajectory information; and the distinguishing module 903 is configured to obtain a distinguishing result by using the visual fusion information and the spatio-temporal modality information, wherein the distinguishing result is used to distinguish each target object in the at least one target object.

[0129] Optionally, the first acquisition module 901 is further configured to acquire distribution information of a plurality of cameras in the target area and a visible range of each camera in the plurality of cameras; divide the target area into a plurality of sub-areas at a target scale based on the distribution information and the visible range; track the at least one target object by using at least one selected camera in each sub-area in the plurality of sub-areas to determine a tracking identifier of each target object in the at least one target object; and acquire the movement trajectory information of the at least one target object in the target area by using the tracking identifier of each target object in the at least one target object.

[0130] Optionally, the second acquisition module 902 is further configured to perform image tracking on the at least one target object to obtain a set of to-be-processed images of each target object in the at least one target object; acquire target appearance features and target shape features of the at least one target object based on the set of to-be-processed images; and acquire the visual fusion information of the at least one target object by using the target appearance features and the target shape features.

[0131] Optionally, the second acquisition module 902 is further configured to perform feature extraction on the to-be-processed images to obtain appearance features and shape features corresponding to each trajectory included in the movement trajectory information; calculate the target appearance features by using the appearance features corresponding to each trajectory, and calculate the target shape features by using the shape features corresponding to each trajectory.

[0132] Optionally, the second obtaining module 902 is further configured to: determine a first similarity matrix between each trajectory included in the movement trajectory information by using the target appearance feature, and determine a second similarity matrix between each trajectory included in the movement trajectory information by using the target shape feature, wherein the first similarity matrix is an appearance similarity matrix, and the second similarity matrix is a shape similarity matrix; and calculate a first mask based on the first similarity matrix and the second similarity matrix, wherein the first mask is a visual fusion mask; and calculate visual fusion information based on the first similarity matrix and the first mask, wherein the visual fusion information is used to represent a visual fusion similarity of the at least one target object.

[0133] Optionally, the second obtaining module 902 is further configured to: generate a second mask and set an initial value of the second mask according to a number of trajectories included in the movement trajectory information, wherein the second mask is a space-time mask; and adjust the initial value based on a space-time constraint relationship between any two trajectories included in the movement trajectory information to obtain a target value of the second mask, so as to determine space-time modal information, wherein the space-time modal information is used to represent a space-time modal feature of the at least one target object.

[0134] Optionally, the distinguishing module 903 is further configured to: calculate a target fusion similarity based on the visual fusion information and the space-time modal information, wherein the target fusion similarity is a multi-modal feature fusion similarity of the at least one target object; and determine the distinguishing result based on the target fusion similarity.

[0135] It should be noted that the first obtaining module 901, the second obtaining module 902, and the distinguishing module 903 correspond to steps S202 to S206 in Embodiment 1, and the three modules have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in Embodiment 1. It should be noted that the modules as part of the device can run in the computer terminal 10 provided in Embodiment 1.

[0136] In the embodiments of the present application, the movement trajectory information of the at least one target object in the target region is obtained, wherein the target region is used to determine the activity range of the at least one target object, and the visual fusion information and the space-time modal information of the at least one target object are obtained based on the movement trajectory information, and the distinguishing result is obtained by using the visual fusion information and the space-time modal information, wherein the distinguishing result is used to distinguish each target object in the at least one target object, so as to achieve the purpose of obtaining the object distinguishing result by using the multi-modal feature fusion information, thereby realizing the technical effect of improving the accuracy of the object distinguishing result, and further solving the technical problem of low accuracy of the object distinguishing result in the processing method of directly distinguishing the object based on the collected features in the related art.

[0137] It should be noted that the preferred embodiments of the present embodiment can refer to the related description in Embodiment 1, which will not be repeated here.

[0138] Embodiment 3

[0139] According to the embodiments of the present application, an embodiment of an electronic device is also provided, which can be any one of the computing devices in the computing device group. The electronic device comprises a processor and a memory, wherein:

[0140] The memory is connected with the processor, and is configured to provide the processor with instructions for processing the following processing steps: obtaining movement trajectory information of at least one target object in a target area, wherein the target area is used to determine the activity range of the at least one target object; obtaining visual fusion information and spatiotemporal modality information of the at least one target object based on the movement trajectory information; and obtaining a distinguishing result by using the visual fusion information and the spatiotemporal modality information, wherein the distinguishing result is used to distinguish each target object in the at least one target object.

[0141] In the embodiments of the present application, by obtaining movement trajectory information of at least one target object in a target area, wherein the target area is used to determine the activity range of the at least one target object, and obtaining visual fusion information and spatiotemporal modality information of the at least one target object based on the movement trajectory information, a distinguishing result is obtained by using the visual fusion information and the spatiotemporal modality information, wherein the distinguishing result is used to distinguish each target object in the at least one target object, so as to achieve the purpose of obtaining an object distinguishing result by multi-modal feature fusion information, thereby realizing the technical effect of improving the accuracy of the object distinguishing result, and further solving the technical problem of low accuracy of the distinguishing result in the related art by directly performing object distinguishing based on the collected features.

[0142] It should be noted that the preferred embodiments of the present embodiment can refer to the related description in Embodiment 1, which will not be repeated here.

[0143] Embodiment 4

[0144] The embodiments of the present application can provide a computer terminal, which can be any one of the computer terminal devices in the computer terminal group. Alternatively, in the present embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.

[0145] Alternatively, in the present embodiment, the computer terminal can be located in at least one network device of a plurality of network devices of a computer network.

[0146] In the embodiment, the computer terminal can execute program codes of the following steps in the method for distinguishing target objects: obtaining movement trajectory information of at least one target object in a target area, wherein the target area is used to determine an activity range of the at least one target object; obtaining visual fusion information and space-time modal information of the at least one target object based on the movement trajectory information; and obtaining a distinguishing result by using the visual fusion information and the space-time modal information, wherein the distinguishing result is used to distinguish each target object in the at least one target object.

[0147] Optionally, Figure 10 is a structural block diagram of another computer terminal according to an embodiment of the present application, as shown in the figure, the computer terminal can include one or more (only one is shown in the figure) processors 122, a memory 124, and a peripheral interface 126. Figure 10

[0148] The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the method and device for distinguishing target objects in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the method for distinguishing target objects described above. The memory can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, and these remote memories can be connected to the computer terminal through a network. Examples of the network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0149] The processor can call information and application programs stored in the memory through the transmission device to execute the following steps: obtaining movement trajectory information of at least one target object in a target area, wherein the target area is used to determine an activity range of the at least one target object; obtaining visual fusion information and space-time modal information of the at least one target object based on the movement trajectory information; and obtaining a distinguishing result by using the visual fusion information and the space-time modal information, wherein the distinguishing result is used to distinguish each target object in the at least one target object.

[0150] ​Optionally, the processor can further execute program codes of the following steps: obtaining distribution information of the plurality of cameras in the target area and a visual range of each camera in the plurality of cameras; dividing the target area at a target scale based on the distribution information and the visual range to obtain a plurality of sub-areas; for each sub-area in the plurality of sub-areas, tracking at least one target object using at least one camera selected in each sub-area to determine a tracking identifier of each target object in the at least one target object; and obtaining movement trajectory information of the at least one target object in the target area using the tracking identifier of each target object in the at least one target object.

[0151] Optionally, the processor can further execute program codes of the following steps: performing image tracking on the at least one target object to obtain a set of to-be-processed images of each target object in the at least one target object; obtaining target appearance features and target body features of the at least one target object based on the set of to-be-processed images; and obtaining visual fusion information of the at least one target object using the target appearance features and the target body features.

[0152] Optionally, the processor can further execute program codes of the following steps: performing feature extraction on the to-be-processed images to obtain appearance features and body features corresponding to each trajectory included in the movement trajectory information; calculating the target appearance features using the appearance features corresponding to each trajectory; and calculating the target body features using the body features corresponding to each trajectory.

[0153] Optionally, the processor can further execute program codes of the following steps: determining a first similarity matrix between each trajectory included in the movement trajectory information using the target appearance features, and determining a second similarity matrix between each trajectory included in the movement trajectory information using the target body features, wherein the first similarity matrix is an appearance similarity matrix, and the second similarity matrix is a body similarity matrix; calculating a first mask based on the first similarity matrix and the second similarity matrix, wherein the first mask is a visual fusion mask; and calculating the visual fusion information using the first similarity matrix and the first mask, wherein the visual fusion information is used to represent a visual fusion similarity of the at least one target object.

[0154] Optionally, the processor can further execute program codes of the following steps: generating a second mask and setting an initial value of the second mask according to a number of trajectories included in the movement trajectory information, wherein the second mask is a space-time mask; adjusting the initial value based on a space-time constraint relationship between any two trajectories included in the movement trajectory information to obtain a target value of the second mask, so as to determine space-time modal information, wherein the space-time modal information is used to represent a space-time modal feature of the at least one target object.

[0155] Optionally, the processor can further execute program codes of the following steps: calculating the target fusion similarity by using the visual fusion information and the space-time modal information, wherein the target fusion similarity is a multi-modal feature fusion similarity of the at least one target object; and determining the distinguishing result by using the target fusion similarity.

[0156] The processor can call the information and the application program stored in the memory by the transmission device to execute the following steps: receiving the moving track information of the at least one target object in the target area from the client, wherein the target area is used to determine the activity range of the at least one target object; obtaining the visual fusion information and the space-time modal information of the at least one target object based on the moving track information, and obtaining the distinguishing result by using the visual fusion information and the space-time modal information, wherein the distinguishing result is used to distinguish each target object in the at least one target object; and feeding back the distinguishing result to the client.

[0157] In the embodiment of the present application, the moving track information of the at least one target object in the target area is obtained, wherein the target area is used to determine the activity range of the at least one target object, and the visual fusion information and the space-time modal information of the at least one target object are obtained based on the moving track information, and the distinguishing result is obtained by using the visual fusion information and the space-time modal information, wherein the distinguishing result is used to distinguish each target object in the at least one target object, so that the object distinguishing result is obtained by using the multi-modal feature fusion information, thereby achieving the technical effect of improving the accuracy of the object distinguishing result, and further solving the technical problem of low accuracy of the object distinguishing result in the related art by directly using the collected features to distinguish the object.

[0158] Those skilled in the art can understand that Figure 10 The structure shown is only schematic, and the computer terminal can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, and the like. Figure 10 The structure of the electronic device is not limited. For example, the computer terminal can further include more or less components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 10 The structure of the electronic device is not limited. For example, the computer terminal can further include more or less components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure. Figure 10 The structure of the electronic device is not limited. For example, the computer terminal can further include more or less components (such as a network interface, a display device, etc.) than those shown in the figure, or have a different configuration from that shown in the figure.

[0159] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiments can be completed by instructing the terminal device related hardware through a program, and the program can be stored in a computer readable storage medium, which can include a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0160] According to the embodiments of the present application, an embodiment of a storage medium is also provided. Optionally, in the present embodiment, the above-mentioned storage medium can be used to save the program code executed by the method for distinguishing target objects provided in the above-mentioned embodiment 1.

[0161] Optionally, in the present embodiment, the above-mentioned storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0162] Optionally, in the present embodiment, the storage medium is configured to store program code for performing the following steps: obtaining movement trajectory information of at least one target object in a target area, wherein the target area is used to determine the activity range of the at least one target object; obtaining visual fusion information and spatio-temporal modal information of the at least one target object based on the movement trajectory information; and obtaining a distinguishing result by using the visual fusion information and the spatio-temporal modal information, wherein the distinguishing result is used to distinguish each target object in the at least one target object.

[0163] Optionally, in the present embodiment, the storage medium is configured to store program code for performing the following steps: obtaining distribution information of a plurality of cameras in a target area and a visible range of each camera in the plurality of cameras; dividing the target area in a target scale based on the distribution information and the visible range to obtain a plurality of sub-areas; for each sub-area in the plurality of sub-areas, tracking at least one target object by using at least one camera selected in each sub-area to determine a tracking identifier of each target object in the at least one target object; and obtaining movement trajectory information of the at least one target object in the target area by using the tracking identifier of each target object in the at least one target object.

[0164] Optionally, in the present embodiment, the storage medium is configured to store program code for performing the following steps: performing image tracking on the at least one target object to obtain a set of to-be-processed images of each target object in the at least one target object; obtaining target appearance features and target shape features of the at least one target object based on the set of to-be-processed images; and obtaining visual fusion information of the at least one target object by using the target appearance features and the target shape features.

[0165] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: performing feature extraction on the to-be-processed image to obtain appearance features and body features corresponding to each trajectory included in the movement trajectory information; calculating target appearance features by using the appearance features corresponding to each trajectory, and calculating target body features by using the body features corresponding to each trajectory.

[0166] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: determining a first similarity matrix between each trajectory included in the movement trajectory information by using the target appearance features, and determining a second similarity matrix between each trajectory included in the movement trajectory information by using the target body features, wherein the first similarity matrix is an appearance similarity matrix, and the second similarity matrix is a body similarity matrix; calculating a first mask based on the first similarity matrix and the second similarity matrix, wherein the first mask is a visual fusion mask; and calculating visual fusion information by using the first similarity matrix and the first mask, wherein the visual fusion information is used to represent visual fusion similarity of the at least one target object.

[0167] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: generating a second mask and setting an initial value of the second mask according to a number of trajectories included in the movement trajectory information, wherein the second mask is a space-time mask; adjusting the initial value based on a space-time constraint relationship between any two trajectories included in the movement trajectory information to obtain a target value of the second mask, so as to determine space-time modal information, wherein the space-time modal information is used to represent space-time modal features of the at least one target object.

[0168] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: calculating target fusion similarity by using the visual fusion information and the space-time modal information, wherein the target fusion similarity is multi-modal feature fusion similarity of the at least one target object; and determining a distinguishing result by using the target fusion similarity.

[0169] Optionally, in the embodiment, the storage medium is configured to store program code for performing the following steps: receiving movement trajectory information of at least one target object in a target area from a client, wherein the target area is used to determine an activity range of the at least one target object; obtaining visual fusion information and space-time modal information of the at least one target object based on the movement trajectory information, and obtaining a distinguishing result by using the visual fusion information and the space-time modal information, wherein the distinguishing result is used to distinguish each target object in the at least one target object; and feeding back the distinguishing result to the client.

[0170] The above-mentioned serial numbers of the embodiments of the application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0171] In the above-described embodiments of the present application, the description of each embodiment focuses on different aspects, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0172] In several embodiments provided by the present application, it should be understood that the disclosed technology can be implemented in other ways. The above-described device embodiments are only illustrative, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, units or modules, and can be electrical or other forms.

[0173] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0174] In addition, each functional unit in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0175] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server or a network device, etc.) execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various program code storage media.

[0176] The above merely describes the preferred embodiments of the present application, and it should be pointed out that, for those skilled in the art, several improvements and refinements can be made without departing from the principles of the present application, and these improvements and refinements should also be considered as falling within the protection scope of the present application.

Claims

1. A method of distinguishing a target object, characterized by, The method comprises the following steps: obtaining movement trajectory information of at least one target object in a target area, wherein the target area is used to determine the activity range of the at least one target object; based on the movement trajectory information, obtaining visual fusion information of the at least one target object; generating a second mask according to the number of trajectories contained in the movement trajectory information and setting an initial value of the second mask, wherein the second mask is a space-time mask; based on the space-time constraint relationship between any two trajectories contained in the movement trajectory information, adjusting the initial value to obtain a target value of the second mask to determine space-time modal information, wherein the space-time modal information is used to represent the space-time modal characteristics of the at least one target object; obtaining a distinction result using the visual fusion information and the space-time modal information, wherein the distinction result is used to distinguish each target object in the at least one target object.

2. The method of claim 1, wherein, Obtaining the movement trajectory information of the at least one target object in the target area comprises: obtaining distribution information of a plurality of cameras in the target area and a visible range of each camera in the plurality of cameras; based on the distribution information and the visible range, dividing the target area at a target scale to obtain a plurality of sub-regions; for each sub-region in the plurality of sub-regions, tracking the at least one target object using at least one camera selected in each sub-region to determine a tracking identifier of each target object in the at least one target object; obtaining the movement trajectory information of the at least one target object in the target area using the tracking identifier of each target object in the at least one target object.

3. The method of claim 1, wherein, Based on the movement trajectory information, obtaining the visual fusion information of the at least one target object comprises: performing image tracking on the at least one target object to obtain a set of to-be-processed images of each target object in the at least one target object; based on the set of to-be-processed images, obtaining target appearance features and target shape features of the at least one target object; using the target appearance features and the target shape features, obtaining the visual fusion information of the at least one target object.

4. The method of claim 3, wherein, Based on the set of to-be-processed images, obtaining the target appearance features and the target shape features of the at least one target object comprises: performing feature extraction on the set of to-be-processed images to obtain appearance features and shape features corresponding to each trajectory contained in the movement trajectory information; using the appearance features corresponding to each trajectory to calculate the target appearance features, and using the shape features corresponding to each trajectory to calculate the target shape features.

5. The method of claim 3, wherein, Using the target appearance features and the target shape features, obtaining the visual fusion information of the at least one target object comprises: determine a first similarity matrix between each trajectory contained in the movement trajectory information by using the target appearance feature, and determine a second similarity matrix between each trajectory contained in the movement trajectory information by using the target body feature, wherein the first similarity matrix is an appearance similarity matrix, and the second similarity matrix is a body similarity matrix; calculate a first mask based on the first similarity matrix and the second similarity matrix, wherein the first mask is a visual fusion mask; calculate the visual fusion information by using the first similarity matrix and the first mask, wherein the visual fusion information is used to represent a visual fusion similarity of the at least one target object.

6. The method of claim 1, wherein, obtaining the distinguishing result by using the visual fusion information and the spatio-temporal modality information includes: calculating a target fusion similarity by using the visual fusion information and the spatio-temporal modality information, wherein the target fusion similarity is a multi-modal feature fusion similarity of the at least one target object; determining the distinguishing result by using the target fusion similarity.

7. A method of distinguishing a target object, characterized by, includes: receiving movement trajectory information of at least one target object in a target area from a client, wherein the target area is used to determine an activity range of the at least one target object; obtaining visual fusion information of the at least one target object based on the movement trajectory information; generating a second mask and setting an initial value of the second mask according to a number of trajectories contained in the movement trajectory information, wherein the second mask is a spatio-temporal mask; adjusting the initial value based on a spatio-temporal constraint relationship between any two trajectories contained in the movement trajectory information to obtain a target value of the second mask, so as to determine spatio-temporal modality information, wherein the spatio-temporal modality information is used to represent spatio-temporal modality features of the at least one target object; obtaining a distinguishing result by using the visual fusion information and the spatio-temporal modality information, wherein the distinguishing result is used to distinguish each target object in the at least one target object; feeding back the distinguishing result to the client.

8. An apparatus for distinguishing a target object, characterized by, includes: a first obtaining module, configured to obtain movement trajectory information of at least one target object in a target area, wherein the target area is used to determine an activity range of the at least one target object; a second obtaining module, configured to obtain visual fusion information of the at least one target object based on the movement trajectory information; generate a second mask and set an initial value of the second mask according to a number of trajectories contained in the movement trajectory information, wherein the second mask is a spatio-temporal mask; adjust the initial value based on a spatio-temporal constraint relationship between any two trajectories contained in the movement trajectory information to obtain a target value of the second mask, so as to determine spatio-temporal modality information, wherein the spatio-temporal modality information is used to represent spatio-temporal modality features of the at least one target object; a distinguishing module, configured to obtain a distinguishing result by using the visual fusion information and the spatio-temporal modality information, wherein the distinguishing result is used to distinguish each target object in the at least one target object.

9. A storage medium, characterized by The storage medium includes a stored program, wherein the program controls a device in which the storage medium is located to execute the method for distinguishing target objects according to any one of claims 1 to 7 when the program is running.

10. A system for differentiating target objects, characterized by Comprise: A processor; And A memory connected with the processor, for providing the processor with instructions for processing the following processing steps: Step 1, obtaining the moving track information of at least one target object in a target area, wherein the target area is used to determine the activity range of the at least one target object; Step 2, based on the moving track information, obtaining the visual fusion information of the at least one target object; according to the number of tracks contained in the moving track information, a second mask is generated and the initial value of the second mask is set, wherein the second mask is a space-time mask; based on the space-time constraint relationship between any two tracks contained in the moving track information, the initial value is adjusted to obtain the target value of the second mask to determine the space-time modal information, wherein the space-time modal information is used to represent the space-time modal characteristics of the at least one target object; Step 3, using the visual fusion information and the space-time modal information to obtain the distinguishing result, wherein the distinguishing result is used to distinguish each target object in the at least one target object.

Citation Information

Patent Citations

  • Pedestrian re-identification method and system based on visual features and space-time constraints

    CN110110598A

  • User behavior data processing method and device

    CN113362090A

  • Clustering method and device, equipment and computer storage medium

    CN113962326A

  • Method, device and system for acquiring moving track information of target object

    CN114091630A

  • Multi-target tracking method for three-dimensional space information fusion based on shielding compensation

    CN114119671A