Information Sending Method, Device, Electronic Device and Computer Readable Medium
By preprocessing and frame reorganizing the multi-directional video sets, a panoramic image sequence is generated, which solves the problem of artificially obtaining target images inaccurately and inefficiently in the prior art, and achieves fast and accurate acquisition of target objects.
Patent Information
- Application Number
- CN202410770415.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-14
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-06-14
AI Technical Summary
In the prior art, it is not accurate enough to obtain the target picture by artificially utilizing image processing software and is inefficient.
By acquiring a multi-directional video set, video preprocessing and frame reorganization are performed, a panoramic image sequence is generated, and packaged and stored with the video frame time series, and the screen information of the target object is extracted from it in response to the acquisition request.
It realizes the quick and accurate acquisition of the screen information of the target object from the shooting video, and meets the predetermined acquisition requirements.
Smart Images

Figure CN118781556B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technologies, and in particular, to an information sending method, apparatus, electronic device, and computer-readable medium. Background Art
[0002] Currently, in the field of security, with the wide application of captured videos in people's daily lives, security awareness has become increasingly important. Object behavior tracking of target objects based on captured videos has become an important development direction in some industries. For determining a target frame of a target object in a captured video, the commonly used method is to manually screen out the target frame corresponding to the target object from the captured video by using image processing software.
[0003] However, when using the above method, the following technical problems often exist:
[0004] The target frame obtained by manually using image processing software to obtain the target frame is often not accurate enough and the acquisition efficiency is low.
[0005] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to ordinary skilled artisans in the country. Summary of the Invention
[0006] The content part of the present disclosure is used to introduce the inventive concept in a brief form, and these inventive concepts will be described in detail in the following detailed implementation part. The content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.
[0007] Some embodiments of the present disclosure propose an information sending method, apparatus, electronic device, and computer-readable medium to solve one or more of the technical problems mentioned in the above background art section.
[0008] In a first aspect, some embodiments of the present disclosure provide an information sending method, including: in response to receiving a security information acquisition request for a target object sent by a security system terminal, acquiring a set of captured videos for the target object, wherein each captured video in the set of captured videos is captured for each shooting direction; performing video preprocessing on each captured video in the set of captured videos to generate a set of captured video frame sequences; performing video frame recombination on the set of captured video frame sequences according to the video frame time corresponding to each captured video frame to obtain a set sequence of captured video frames, wherein the video frame times corresponding to the captured video frames in each set of captured video frames are the same; generating a panoramic captured image sequence for the set sequence of captured video frames; corresponding and packaging the panoramic captured image sequence and the corresponding video frame time sequence for storage to obtain a first image packaging file, wherein there is a time correspondence relationship between the video frame time sequence and the set sequence of captured video frames; in response to receiving object picture segment acquisition information for the target object, acquiring at least one panoramic captured image and the corresponding at least one video frame time for the object picture segment acquisition information from the stored first image packaging file, wherein the object picture segment acquisition information includes: object picture acquisition direction information and object picture acquisition time; determining at least one set of captured video frames corresponding to the at least one video frame time; and sending the at least one panoramic captured image, the at least one video frame time, and the at least one set of captured video frames to the security system terminal.
[0009] Second aspect, some embodiments of the present disclosure provide an information sending device, including: a first obtaining unit configured to obtain a set of captured videos for the target object in response to receiving a security information obtaining request for the target object sent by a security system terminal, where each captured video in the set of captured videos is captured for each shooting direction; a video preprocessing unit configured to perform video preprocessing on each captured video in the set of captured videos to generate a set of captured video frame sequences; a recombination unit configured to perform video frame recombination on the set of captured video frame sequences according to the video frame time corresponding to each captured video frame to obtain a set sequence of captured video frames, where the video frame times corresponding to the captured video frames in each set of captured video frames are the same; a generating unit configured to generate a panoramic captured image sequence for the set sequence of captured video frames; a packaging and storing unit configured to perform corresponding packaging and storing on the panoramic captured image sequence and the corresponding video frame time sequence to obtain a first image packaging file, where there is a time correspondence relationship between the video frame time sequence and the set sequence of captured video frames; a second obtaining unit configured to obtain at least one panoramic captured image and the corresponding at least one video frame time for the object picture segment obtaining information from the stored first image packaging file in response to receiving the object picture segment obtaining information for the target object, where the object picture segment obtaining information includes: object picture obtaining direction information and object picture obtaining time; a determining unit configured to determine at least one set of captured video frames corresponding to the at least one video frame time; a sending unit configured to send the at least one panoramic captured image, the at least one video frame time, and the at least one set of captured video frames to the security system terminal.
[0010] Third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any implementation manner of the first aspect.
[0011] Fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, where when the program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.
[0012] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the information sending method of some embodiments of the present disclosure, the picture information corresponding to the target object under the predetermined acquisition requirements can be obtained quickly and accurately from the captured video set. Specifically, the reasons for the lack of accuracy and efficiency of the relevant picture information are as follows: When obtaining the target picture by artificially using image processing software, the obtained target picture is often inaccurate and the acquisition efficiency is low. Based on this, in the information sending method of some embodiments of the present disclosure, first, in response to receiving a security information acquisition request for a target object sent by a security system terminal, a captured video set for the target object is obtained, where each captured video in the captured video set is captured for each shooting direction. Here, through the multi-directional captured videos corresponding to the target object, an all-round video of the target object can be obtained, ensuring that the target picture under the preset acquisition requirements can be obtained subsequently. Then, video preprocessing is performed on each captured video in the captured video set to generate a captured video frame sequence set, so as to facilitate the subsequent acquisition of the corresponding target picture. Next, according to the video frame time corresponding to each captured video frame, the captured video frame sequence set is reorganized to obtain a captured video frame set sequence, so as to facilitate the subsequent generation of a panoramic captured image. Among them, the video frame times corresponding to the captured video frames in each captured video frame set are the same. Then, a panoramic captured image sequence for the captured video frame set sequence can be accurately generated. Here, by generating the corresponding panoramic captured image, subsequent information is obtained based on the object picture segment to quickly and accurately determine the corresponding target picture (that is, a certain captured video frame in the captured video set). Immediately afterwards, the panoramic captured image sequence and the corresponding video frame time sequence are packaged and stored correspondingly to obtain a first image package file. Among them, the video frame time sequence has a time correspondence relationship with the captured video frame set sequence. Here, through the first image package file, it is convenient to obtain the target picture for information acquisition of the object picture segment in real time subsequently. Secondly, in response to receiving the object picture segment acquisition information (i.e., the predetermined acquisition requirements) for the target object, at least one panoramic captured image and the corresponding at least one video frame time for the object picture segment acquisition information are quickly and accurately obtained from the stored first image package file. Among them, the object picture segment acquisition information includes: object picture acquisition direction information and object picture acquisition time. Then, at least one captured video frame set corresponding to the at least one video frame time is determined. Finally, the at least one panoramic captured image, the at least one video frame time, and the at least one captured video frame set are sent to the security system terminal, so as to obtain the target picture corresponding to the target object with the predetermined acquisition requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic and the elements and elements are not necessarily drawn to scale.
[0014] Figure 1 is a flowchart of some embodiments of an information sending method according to the present disclosure;
[0015] Figure 2 is a schematic structural diagram of some embodiments of an information sending device according to the present disclosure;
[0016] Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Specific Embodiments
[0017] The embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0018] In addition, it should be noted that for the sake of convenience of description, only the parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0019] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0020] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0022] The present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.
[0023] Reference Figure 1, which shows the process 100 of some embodiments of the information sending method according to the present disclosure. The information sending method includes the following steps:
[0024] Step 101, obtain a set of captured videos for a target object.
[0025] In some embodiments, in response to receiving a security information acquisition request for a target object sent by a security system terminal, the execution entity of the above information sending method can obtain a set of captured videos for the target object through a wired connection method or a wireless connection method. In practice, the target object can be a target person in a security scenario or a target moving vehicle in a security parking lot. Each captured video in the set of captured videos is captured for each shooting direction. For example, each shooting direction can include: the front direction, the rear direction, the left direction, and the right direction. Among them, the security system terminal is a terminal equipped with a security system. The security system can include: a video monitoring system. The video monitoring system can monitor and record the on-site situation in real time, providing intuitive and reliable on-site information. The video monitoring system usually consists of cameras, transmission devices, storage devices, and display devices, etc. As the core device of the video monitoring system, the camera can capture and record image information. According to different application scenarios, the camera can be divided into indoor and outdoor categories, and can provide options of high definition, standard definition, and ordinary image resolution. The transmission device is responsible for transmitting the image information captured by the camera to the storage device. Common transmission methods include fiber optic transmission, network cable transmission, and wireless transmission, etc. The storage device is used to save the image information transmitted by the camera. Currently, common storage methods include digital video recorders and cloud storage. The display device is used to display the image information transmitted by the camera, and usually includes a video wall, a computer monitor, etc.
[0026] Step 102, perform video preprocessing on each captured video in the set of captured videos to generate a set of captured video frame sequences.
[0027] In some embodiments, the above execution entity can perform video preprocessing on each captured video in the set of captured videos to generate a set of captured video frame sequences. Among them, there is a one-to-one correspondence between the captured video frame sequences in the set of captured video frame sequences and the captured videos in the set of captured videos. In practice, the video preprocessing can include but is not limited to at least one of the following: video decoding processing, video frame extraction processing.
[0028] Step 103, perform video frame recombination on the set of captured video frame sequences according to the video frame time corresponding to each captured video frame to obtain a set of captured video frame sequences.
[0029] In some embodiments, the above-mentioned execution entity may reorganize the set of captured video frame sequences according to the video frame time corresponding to each captured video frame to obtain a sequence of sets of captured video frames. Among them, the video frame times corresponding to the respective captured video frames in each set of captured video frames are the same.
[0030] As an example, the above-mentioned execution entity may determine at least one captured video frame under the same video frame time in the set of captured video frame sequences as a set of captured video frames to obtain a sequence of sets of captured video frames.
[0031] Step 104, generate a sequence of panoramic captured images for the sequence of sets of captured video frames.
[0032] In some embodiments, the above-mentioned execution entity may generate a sequence of panoramic captured images for the sequence of sets of captured video frames. Among them, there is a one-to-one correspondence between the sets of captured video frames in the sequence of sets of captured video frames and the panoramic captured images in the sequence of panoramic captured images. The panoramic captured image may be a panoramic image composed of the respective captured video frames in the corresponding set of captured video frames.
[0033] In some optional implementation manners of some embodiments, the generating a sequence of panoramic captured images for the sequence of sets of captured video frames may include the following steps:
[0034] The first step, for the target video frame time in the video frame time sequence, execute the following second generation step:
[0035] The first sub-step, in response to determining that the target video frame time is not the first video frame time in the video frame time sequence, determine the previous video frame time corresponding to the target video frame time. Among them, the previous video frame time may be a video frame time adjacent to the target video frame time and earlier than the target video frame. The first video frame time is the earliest video frame in the video frame time sequence.
[0036] The second sub-step, determine the panoramic captured image corresponding to the previous video frame time as the previous panoramic captured image.
[0037] The third sub-step, input the previous panoramic captured image into a pre-trained image feature extraction model to generate previous panoramic image feature information. Among them, the image feature extraction model may be a neural network model for extracting image semantic features. In practice, the image feature extraction model may be a 15-layer cascaded convolutional neural network model. The previous panoramic image feature information may characterize the image feature semantics of the previous panoramic image.
[0038] The fourth sub-step is to determine the set of captured video frames corresponding to the target video frame time and the corresponding set of captured azimuth information. Among them, the captured azimuth information can be the specific azimuth content of the shooting azimuth. There is a one-to-one correspondence between the captured video frames in the set of captured video frames and the captured azimuth information in the set of captured azimuth information.
[0039] The fifth sub-step is to input each captured video frame in the set of captured video frames into the image feature extraction model to generate captured image feature information, obtaining a set of captured image feature information. Among them, the captured image feature information can characterize the image feature semantics corresponding to the captured image.
[0040] The sixth sub-step is to multiply the previous panoramic image feature information by a preset matrix to generate a multiplication matrix. Among them, the preset matrix can be determined in advance by relevant experts and represents the change information of the panoramic image between successive video frame times. The preset matrix can also be learned by a relevant neural network. The multiplication matrix can indirectly characterize the alternative feature information corresponding to the panoramic image to be generated.
[0041] The seventh sub-step is to input the set of captured image feature information and the multiplication matrix into a pre-trained image feature weight information generation model to generate a set of image feature weight information as the first initial set of image feature weight information. Among them, the first initial image feature weight information generation model can be a neural network model for generating image feature weight information. The first initial image feature weight information can characterize the importance degree or content degree of the object feature information corresponding to the captured image feature information. The first initial image feature weight information can be a value between 0 and 1. The larger the value, the higher the importance degree or content degree of the object feature information corresponding to the captured image feature information. In practice, the image feature weight information generation model can be an attention mechanism model based on a convolutional neural network model. The image feature weight information generation model can be a CoAtNet model based on the attention mechanism.
[0042] The eighth sub-step is to perform encoding processing on each captured azimuth information in the set of captured azimuth information to generate a set of azimuth encoding information.
[0043] As an example, the above-mentioned execution entity can input each captured azimuth information in the set of captured azimuth information into an information encoding model to generate a set of azimuth encoding information. Among them, the azimuth encoding information can be information in vector form representing the azimuth semantics corresponding to the captured azimuth information. The information encoding model can be a pre-trained model.
[0044] The ninth sub-step is to perform corresponding information splicing on the first initial set of image feature weight information, the set of captured image feature information, and the set of azimuth encoding information to generate a first set of spliced information.
[0045] As an example, for each piece of first initial feature weight information in the first initial feature weight information set, the above-mentioned execution entity may splice the first initial feature weight information, the corresponding captured image feature information, and the corresponding orientation coding information to generate first splicing information.
[0046] The tenth sub-step is to input the first splicing information set into the attention mechanism model based on orientation and feature content to generate a first image feature weight information set for the captured image feature information set. Among them, there is a one-to-one correspondence between the captured image feature information in the captured image feature information set and the first image feature weight information in the first image feature weight information set. The first image feature weight information can represent the importance degree or content degree of the object feature information corresponding to the captured image feature information. The first image feature weight information can be a value between 0 and 1. The larger the value, the higher the importance degree or content degree of the object feature information corresponding to the captured image feature information. The attention mechanism model based on orientation and feature content can be a multi-head attention mechanism model.
[0047] The eleventh sub-step is to generate a panoramic captured image corresponding to the target video frame time according to the captured image feature information set and the first image feature weight information set by using a pre-trained generative and adversarial neural network model. Among them, the generative and adversarial neural network model can be a GAN model.
[0048] As an example, first, the above-mentioned execution entity may perform corresponding weighted summation on the captured image feature information set and the first image feature weight information set to generate a weighted summation information set. Then, according to the weighted summation information set, use the generative and adversarial neural network model to generate a panoramic captured image corresponding to the target video frame time.
[0049] The twelfth sub-step is to, in response to determining that the target video frame time is the second video frame time, sort each panoramic captured image in the obtained panoramic video image set to generate a panoramic video image sequence.
[0050] The second step is to, in response to determining that the target video frame time is not the second video frame time, determine the next video frame time corresponding to the target video frame time as the target video frame time, and continue to execute the generation step. The second video frame time is the video frame with the latest time in the video frame time sequence.
[0051] Optionally, generating a panoramic captured image corresponding to the target video frame time according to the captured image feature information set and the first image feature weight information set by using a pre-trained generative and adversarial neural network model includes the following steps:
[0052] First, vectorize each piece of captured image feature information in the captured image feature information set to generate captured image feature vectors, obtaining a captured image feature vector set, and vectorize each piece of second image feature weight information in the second image feature weight information set to generate weight vectors, obtaining a weight vector set.
[0053] Second, sequentially combine the captured image feature vectors in the captured image feature vector set with the corresponding weight vectors in the weight vector set in order to generate combined information, obtaining a combined information set. Among them, the combined information can be [captured image feature vector, weight vector].
[0054] Third, input each piece of combined information in the combined information set into the generative model included in the generative and adversarial neural network model to generate an initial panoramic captured image. Among them, the generative model can be a multi-layer cascaded residual model.
[0055] Fourth, input the initial panoramic captured image into the adversarial model included in the generative and adversarial neural network model to generate an image feature information set. Among them, the adversarial model is a neural network model that generates the corresponding feature content of the captured image feature set. The adversarial model includes: an image feature extraction model and a feature information generation model. The feature information generation model can be a multi-layer cascaded fully connected layer. There is a one-to-one correspondence between the captured images features in the captured image feature set and the captured image feature vectors in the captured image feature vector set. Among them, the adversarial model can be trained periodically
[0056] Fifth, generate a corresponding image feature vector set for the image feature information set, where there is a one-to-one correspondence between the image feature information in the image feature information set and the image feature vectors in the image feature vector set.
[0057] Sixth, determine the vector similarity between each piece of image feature information in the image feature vector set and the corresponding captured image feature vector in the captured image feature vector set, obtaining a vector similarity set. Among them, the vector similarity can characterize the semantic content similarity between the image feature information and the captured image feature vector.
[0058] Seventh, in response to the non-existence of a vector similarity less than the target similarity value in the vector similarity set, determine the initial panoramic captured image as the panoramic captured image.
[0059] In the eighth step, in response to the existence of a vector similarity greater than or equal to the target similarity value in the vector similarity set, select the candidate generative model at the next position corresponding to the generative model from the candidate generative model sequence. Among them, each candidate generative model in the candidate generative model sequence is sorted in turn according to the model accuracy and the model calculation amount. That is, the closer the candidate generative model is to the front position in the candidate generative model sequence, the lower the model accuracy and the smaller the corresponding model calculation amount. Each candidate generative model in the candidate generative model sequence can be a model with various network structures. For example, it may include: a convolutional layer in series with multiple layers, a residual layer in series with multiple layers. Determine the candidate generative model as the generative model included in the generative adversarial neural network model, and perform the generation of the panoramic captured image again.
[0060] In some optional implementation manners of some embodiments, after generating the panoramic captured image corresponding to the target video frame time by using the pre-trained generative adversarial neural network model according to the captured image feature information set and the first image feature weight information set, the method further includes:
[0061] In the first step, input the panoramic captured image into the image feature extraction model to generate panoramic image feature information.
[0062] In the second step, determine the feature difference matrix between the panoramic image feature information and the previous panoramic image feature information. Among them, both the panoramic image feature information and the previous panoramic image feature information can be feature information in matrix form. The feature difference matrix can represent the semantic gap between the panoramic image feature information and the previous panoramic image feature information. Among them, the multiplication matrix between the feature difference matrix and the previous panoramic image feature information is the panoramic image feature information.
[0063] In the third step, determine the first difference information between the feature difference matrix and the preset matrix.
[0064] As an example, the above-mentioned execution subject can subtract the preset matrix from the feature difference matrix to generate a first difference matrix.
[0065] In the fourth step, input the captured image feature information set and the panoramic image feature information into the image feature weight information generation model to generate an actual image feature weight information set as the actual image feature weight information set.
[0066] In the fifth step, determine the second difference information between the actual image feature weight information set and the initial image feature weight information set.
[0067] As an example, the above-mentioned execution entity may subtract the initial image feature weight information in the initial image feature weight information set from the actual image feature weight information in the actual image feature weight information set to generate a subtraction value set as the second difference information.
[0068] Sixth step, input the panoramic captured image and the previous panoramic captured image into a pre-trained image content difference information generation model to generate image content difference information. Among them, the image content difference information generation model is a neural network model for generating image content difference information. In practice, the image content difference information can represent the image content difference between the panoramic captured image and the previous panoramic captured image.
[0069] Seventh step, in response to determining that the first difference information meets the first preset difference condition, the second difference information meets the second preset difference condition, and the image content difference information meets the third preset difference condition, generate verification information indicating that the panoramic captured image passes the verification. The first preset difference condition may be that the element sizes of each element in the matrix corresponding to the first difference information are within a first preset interval. The second preset difference condition may be that the numerical sizes of each subtraction value in the subtraction value set corresponding to the second difference information are within a second preset interval. The third preset difference condition may be that the content difference degree corresponding to the image content difference information is less than a predetermined degree.
[0070] In some optional implementation manners of some embodiments, before determining the previous video frame time corresponding to the target video frame time in response to determining that the target video frame time is not the first video frame time in the video frame time series, the method further includes:
[0071] First step, in response to determining that the target video frame time is the first video frame time in the video frame time series, determine the captured video frame set corresponding to the target video frame time and the corresponding captured orientation information set.
[0072] Second step, use panoramic image generation technology to generate a candidate panoramic captured image for the captured video frame set.
[0073] Third step, input the captured image feature information set and the candidate panoramic captured image into the image feature weight information generation model to generate an image feature weight information set as a candidate image feature weight information set.
[0074] Fourth step, perform corresponding information splicing on the candidate image feature weight information set, the second captured image feature information set, and the orientation coding information set to generate a second splicing information set.
[0075] Step 5: Input the second splicing information set into the attention mechanism model based on orientation and feature content to generate a second image feature weight information set for the captured image feature information set.
[0076] Step 6: According to the captured image feature information set and the second image feature weight information set, use the pre-trained generative and adversarial neural network model to generate a panoramic captured image corresponding to the target video frame time.
[0077] Step 105: Package and store the panoramic captured image sequence and the corresponding video frame time sequence in a corresponding manner to obtain a first image package file.
[0078] In some embodiments, the above-mentioned execution entity may package and store the panoramic captured image sequence and the corresponding video frame time sequence in a corresponding manner to obtain a first image package file. Among them, there is a time correspondence between the video frame time sequence and the captured video frame set sequence. The video frame time in the video frame time sequence has a one-to-one correspondence with the captured video frame set in the captured video frame set sequence.
[0079] As an example, the above-mentioned execution entity may package the panoramic captured image sequence and the corresponding video frame time sequence in the form of key-value pairs.
[0080] In some optional implementation manners of some embodiments, after step 105, the steps further include:
[0081] First step: Generate first file display information corresponding to the first image package file. The first file display information may be file description display information representing the file content corresponding to the first image package file. For example, the first file display information may be "picture-related content information of the target picture corresponding to the target object in the panoramic captured image sequence".
[0082] Second step: Display the first file display information on a display terminal for the security system terminal to obtain the corresponding image content;
[0083] Step 106: In response to receiving the object picture segment acquisition information for the target object, obtain at least one panoramic captured image and the corresponding at least one video frame time for the object picture segment acquisition information from the stored first image package file.
[0084] In some embodiments, in response to receiving the object picture segment acquisition information for the target object, the above-mentioned execution entity may obtain at least one panoramic captured image and the corresponding at least one video frame time for the object picture segment acquisition information from the stored first image package file. The object picture segment acquisition information may include: object picture acquisition orientation information and object picture acquisition time. The object picture acquisition orientation information may be the information for acquiring the picture of an object in a certain orientation. For example, the object picture acquisition orientation information may be "acquire the picture in the orientation directly opposite to the target object". The object picture acquisition time may be the information for acquiring the picture of an object at a certain time. For example, the object picture acquisition time information may be "acquire the picture of the target object at the target time". There is a one-to-one correspondence between the panoramic captured images in the at least one panoramic captured image and the at least one video frame time.
[0085] Step 107, determine at least one set of captured video frames corresponding to the at least one video frame time.
[0086] In some embodiments, the above-mentioned execution entity may determine at least one set of captured video frames corresponding to the at least one video frame time. Among them, there is a one-to-one correspondence between the video frame time in the at least one video frame time and the set of captured video frames in the at least one set of captured video frames.
[0087] Step 108, send the at least one panoramic captured image, the at least one video frame time, and the at least one set of captured video frames to the security system terminal.
[0088] In some embodiments, the above-mentioned execution entity may send the at least one panoramic captured image, the at least one video frame time, and the at least one set of captured video frames to the security system terminal. Among them, the security system terminal may be a terminal for acquiring relevant picture information of the corresponding picture of the target object.
[0089] In some optional implementation manners of some embodiments, after step 108, the steps further include:
[0090] First step, screen out a target number of key captured video frames from the set of captured video frame sequences as the set of key captured video frames. Among them, the key captured video frames may be the video frames with key captured content.
[0091] Second step, determine the video frame time sequence corresponding to the set of key captured video frames. Among them, there is a one-to-one correspondence between the key captured video frames in the set of key captured video frames and the video frame time in the video frame time sequence.
[0092] Third step, for each video frame time in the video frame time sequence, execute the following first generation step:
[0093] Sub-step 1: For each captured video in the captured video set, determine the captured image in the captured video whose corresponding frame time is the video frame time as the target captured image.
[0094] Sub-step 2: Generate a key panoramic captured image according to the obtained target captured image set.
[0095] As an example, the above execution entity can use panoramic image generation technology to generate a key panoramic captured image for the target captured image set.
[0096] Fourth step: Package and store the obtained key panoramic captured image sequence and the corresponding key video frame time sequence to obtain a second image package file.
[0097] Fifth step: Generate second file display information for the second image package file. The second file display information can be file description display information characterizing the file content corresponding to the second image package file. For example, the second file display information can be "picture-related content information of the target picture corresponding to the target object for the key panoramic captured image sequence".
[0098] Sixth step: Display the second file display information on a display terminal for the security system terminal to obtain the corresponding key image content.
[0099] Seventh step: In response to receiving key image content acquisition information for the target object, obtain at least one key panoramic captured image and the corresponding at least one key video frame time from the stored second image package file. Among them, there is a one-to-one correspondence between the key panoramic captured image in the at least one key panoramic captured image and the key video frame time in the at least one key video frame time. The key image content acquisition information can include: key image acquisition information and key image time acquisition information. The key image acquisition information can be information about obtaining the picture of an acquisition object at a certain key video frame. For example, the key image acquisition information can be "obtain the picture corresponding to the key video frame of the target object". The key image time acquisition information can be information about obtaining the picture of an acquisition object at a certain key time. For example, the key image time acquisition information can be "obtain the picture of the target object at the target key time".
[0100] Eighth step: Determine at least one key captured video frame set corresponding to the at least one key video frame time. Among them, there is a one-to-one correspondence between the key video frame time in the at least one key video frame time and the key captured video frame set in the at least one key captured video frame set.
[0101] Step 9: Send the at least one key panoramic captured image, the at least one key video frame time, and the at least one key captured video frame set to the security system terminal.
[0102] Optionally, screening out a target number of key captured video frames from the captured video frame sequence set as the key captured video frame set may include the following steps:
[0103] Step 1: Obtain a video usage information set for the captured video set. Among them, the video usage information may be the usage information of the captured video. For example, the video usage information may be the usage for tracking abnormal behaviors of an object, or the usage for evaluating the work conscientiousness of an object. Each video usage information in the video usage information set may be preset.
[0104] Step 2: For each video usage information in the video usage information set, perform the following third generation step:
[0105] Sub-step 1: Determine at least one attribute label corresponding to the video usage information. Among them, the attribute label may be a label of the usage attribute corresponding to the video usage information. The at least one attribute label may be preset. For example, for the video usage information being the usage for tracking abnormal behaviors of an object, the corresponding at least one attribute label may include: action posture label, object expression label.
[0106] Sub-step 2: Obtain at least one attribute label information generation model corresponding to the at least one attribute label. Among them, there is a one-to-one correspondence between the attribute label in the at least one attribute label and the attribute label information generation model in the at least one attribute label information generation model. The attribute label information generation model is a neural network model that is pre-trained to generate the label content corresponding to the attribute label. For example, the attribute label information generation model may be a multi-layer cascaded convolutional neural network model.
[0107] Sub-step 3: For each attribute label information generation model in the at least one attribute label information generation model, perform the following fourth generation step:
[0108] First sub-step: Input each captured video frame in the captured video frame sequence set into the attribute label information generation model to generate attribute label information, obtaining an attribute label information sequence set, where the attribute label information includes: label name information and label probability information. Among them, the label name information may be the label name corresponding to the attribute label. The label probability information may be the probability that the corresponding captured video frame corresponds to the attribute label being the attribute label information. In practice, the larger the label probability information, the greater the probability that the corresponding captured video frame corresponds to the attribute label being the attribute label information.
[0109] Step 3: Screen out the captured video frames from the captured video frame sequence set, where the corresponding tag name information is the same as the target attribute tag information and the corresponding tag probability information is the largest, as the key captured video frames, where the target attribute tag information is the attribute tag information corresponding to the attribute tag information generation model.
[0110] The above embodiments of the present disclosure have the following beneficial effects: Through the information sending method of some embodiments of the present disclosure, the picture information corresponding to the target object under the predetermined acquisition requirements can be obtained quickly and accurately from the captured video set. Specifically, the reasons for the lack of accuracy and efficiency of the relevant picture information are as follows: By manually using image processing software to obtain the target picture, the obtained target picture is often inaccurate and the acquisition efficiency is low. Based on this, in the information sending method of some embodiments of the present disclosure, first, in response to receiving a security information acquisition request for a target object sent by a security system terminal, a captured video set for the target object is obtained, where each captured video in the captured video set is captured for each shooting direction. Here, through the multi-directional captured videos corresponding to the target object, an all-round video of the target object can be obtained, ensuring that the target picture under the preset acquisition requirements can be obtained subsequently. Then, video preprocessing is performed on each captured video in the captured video set to generate a captured video frame sequence set, so as to facilitate the subsequent acquisition of the corresponding target picture. Next, according to the video frame time corresponding to each captured video frame, the captured video frame sequence set is reorganized to obtain a captured video frame set sequence, so as to facilitate the subsequent generation of a panoramic captured image. Among them, the video frame times corresponding to the captured video frames in each captured video frame set are the same. Then, a panoramic captured image sequence for the captured video frame set sequence can be accurately generated. Here, by generating the corresponding panoramic captured image, subsequent information is obtained based on the object picture segment to quickly and accurately determine the corresponding target picture (i.e., a certain captured video frame in the captured video set). Immediately afterwards, the panoramic captured image sequence and the corresponding video frame time sequence are packaged and stored correspondingly to obtain a first image package file. Among them, the video frame time sequence has a time correspondence relationship with the captured video frame set sequence. Here, through the first image package file, it is convenient to obtain the target picture for subsequent real-time acquisition of information based on the object picture segment. Secondly, in response to receiving the object picture segment acquisition information (i.e., the predetermined acquisition requirements) for the target object, at least one panoramic captured image and the corresponding at least one video frame time for the object picture segment acquisition information are quickly and accurately obtained from the stored first image package file. Among them, the object picture segment acquisition information includes: object picture acquisition direction information and object picture acquisition time. Then, at least one captured video frame set corresponding to the at least one video frame time is determined. Finally, the at least one panoramic captured image, the at least one video frame time, and the at least one captured video frame set are sent to the security system terminal, so as to obtain the target picture corresponding to the target object with the predetermined acquisition requirements.
[0111] Further reference Figure 2, as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of an information sending device, and these device embodiments correspond to Figure 1 the method embodiments shown, and the information sending device can be specifically applied to various electronic devices.
[0112] As shown in Figure 2 , an information sending device 200 includes: a first acquisition unit 201, a video preprocessing unit 202, a recombination unit 203, a generation unit 204, a packaging and storage unit 205, a second acquisition unit 206, a determination unit 207, and a sending unit 208. Among them, the first acquisition unit 201 is configured to, in response to receiving a security information acquisition request for a target object sent by a security system terminal, acquire a set of captured videos for the target object, where each captured video in the set of captured videos is captured for each shooting direction; the video preprocessing unit 202 is configured to perform video preprocessing on each captured video in the set of captured videos to generate a set of captured video frame sequences; the recombination unit 203 is configured to perform video frame recombination on the set of captured video frame sequences according to the video frame time corresponding to each captured video frame to obtain a set of captured video frame sequences, where the video frame times corresponding to the captured video frames in each set of captured video frames are the same; the generation unit 204 is configured to generate a panoramic captured image sequence for the set of captured video frame sequences; the packaging and storage unit 205 is configured to perform corresponding packaging and storage on the panoramic captured image sequence and the corresponding video frame time sequence to obtain a first image packaging file, where the video frame time sequence has a time correspondence relationship with the set of captured video frame sequences; the second acquisition unit 206 is configured to, in response to receiving object picture segment acquisition information for the target object, acquire at least one panoramic captured image and the corresponding at least one video frame time for the object picture segment acquisition information from the stored first image packaging file, where the object picture segment acquisition information includes: object picture acquisition direction information and object picture acquisition time; the determination unit 207 is configured to determine at least one set of captured video frames corresponding to the at least one video frame time; the sending unit 208 is configured to send the at least one panoramic captured image, the at least one video frame time, and the at least one set of captured video frames to the security system terminal.
[0113] It can be understood that the units described in the information sending device 200 correspond to the respective steps in the method described with reference to Figure 1 . Therefore, the operations, features, and beneficial effects described above for the method also apply to the information sending device 200 and the units included therein, and will not be repeated here.
[0114] Next, with reference to Figure 3, which shows a schematic structural diagram of an electronic device (e.g., an electronic device) 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The illustrated electronic device is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.
[0115] As Figure 3 shown, the electronic device 300 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 301, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0116] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 the electronic device 300 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be alternatively implemented or included. Figure 3 Each block shown in
[0117] Specifically, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from a network via the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above-mentioned functions defined in the methods of some embodiments of the present disclosure are performed.
[0118] It should be noted that in some embodiments of the present disclosure, the above-mentioned computer-readable medium may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0119] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0120] The above computer-readable medium may be included in the above electronic device; or it may exist independently without being assembled into the electronic device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device is caused to: in response to receiving a security information acquisition request for a target object sent by a security system terminal, acquire a set of captured videos for the target object, wherein each captured video in the set of captured videos is captured for each shooting orientation; perform video preprocessing on each captured video in the set of captured videos to generate a set of captured video frame sequences; perform video frame recombination on the set of captured video frame sequences according to the video frame time corresponding to each captured video frame to obtain a sequence of sets of captured video frames, wherein the video frame times corresponding to the captured video frames in each set of captured video frames are the same; generate a panoramic captured image sequence for the sequence of sets of captured video frames; perform corresponding packaging and storage on the panoramic captured image sequence and the corresponding video frame time sequence to obtain a first image packaging file, wherein there is a time correspondence relationship between the video frame time sequence and the sequence of sets of captured video frames; in response to receiving object picture segment acquisition information for the target object, acquire at least one panoramic captured image and the corresponding at least one video frame time for the object picture segment acquisition information from the stored first image packaging file, wherein the object picture segment acquisition information includes: object picture acquisition orientation information and object picture acquisition time; determine at least one set of captured video frames corresponding to the at least one video frame time; and send the at least one panoramic captured image, the at least one video frame time, and the at least one set of captured video frames to the security system terminal.
[0121] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages - such as Java, Smalltalk, C++; and also include conventional procedural programming languages - such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or may be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0123] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes a first acquisition unit, a video preprocessing unit, a recombination unit, a generation unit, a packaging and storage unit, a second acquisition unit, a determination unit, and a transmission unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring a set of captured videos of a target object".
[0124] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0125] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, technical solutions formed by mutually replacing the above features with (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.
Claims
1. An information sending method, comprising: Upon receiving a security information acquisition request for a target object sent by a security system terminal, acquiring a set of captured videos for the target object, wherein each captured video in the set of captured videos is captured for each shooting direction; Performing video preprocessing on each captured video in the set of captured videos to generate a set of captured video frame sequences; Recombining the set of captured video frame sequences according to the video frame time corresponding to each captured video frame to obtain a set sequence of captured video frames, wherein the video frame times corresponding to the captured video frames in each set of captured video frames are the same; Generating a panoramic captured image sequence for the set sequence of captured video frames, wherein generating the panoramic captured image sequence for the set sequence of captured video frames includes: for a target video frame time in the video frame time sequence, performing the following second generation step: upon determining that the target video frame time is not the first video frame time in the video frame time sequence, determining the previous video frame time corresponding to the target video frame time; determining the panoramic captured image corresponding to the previous video frame time as the previous panoramic image; inputting the previous panoramic image into a pre-trained image feature extraction model to generate previous panoramic image feature information; determining the set of captured video frames and the corresponding set of shooting direction information corresponding to the target video frame time; inputting each captured video frame in the set of captured video frames into the image feature extraction model to generate captured image feature information, obtaining a set of captured image feature information; multiplying the previous panoramic image feature information by a preset matrix to generate a multiplication matrix; inputting the set of captured image feature information and the multiplication matrix into a pre-trained image feature weight information generation model to generate a set of image feature weight information as a first initial set of image feature weight information; performing encoding processing on each shooting direction information in the set of shooting direction information to generate a set of direction encoding information; performing corresponding information splicing on the first initial set of image feature weight information, the set of captured image feature information, and the set of direction encoding information to generate a first spliced information set; inputting the first spliced information set into an attention mechanism model based on direction and feature content to generate a first set of image feature weight information for the set of captured image feature information; according to the set of captured image feature information and the first set of image feature weight information, using a pre-trained generative and adversarial neural network model to generate a panoramic captured image corresponding to the target video frame time; upon determining that the target video frame time is the second video frame time, sorting the panoramic captured images in the obtained panoramic video image set to generate a panoramic video image sequence; upon determining that the target video frame time is not the second video frame time, determining the next video frame time corresponding to the target video frame time as the target video frame time, and continuing to perform the generation step; Correspondingly package and store the panoramic captured image sequence and the corresponding video frame time sequence to obtain a first image package file, where the video frame time sequence has a time correspondence with the captured video frame set sequence; In response to receiving object picture segment acquisition information for the target object, obtain at least one panoramic captured image and the corresponding at least one video frame time for the object picture segment acquisition information from the stored first image package file, where the object picture segment acquisition information includes: object picture acquisition orientation information and object picture acquisition time; Determine at least one captured video frame set corresponding to the at least one video frame time; Send the at least one panoramic captured image, the at least one video frame time, and the at least one captured video frame set to the security system terminal.
2. The method according to claim 1, wherein The method further includes: Screen a target number of key captured video frames from the captured video frame sequence set as a key captured video frame set; Determine the video frame time sequence corresponding to the key captured video frame set; For each video frame time in the video frame time sequence, perform the following first generation step: For each captured video in the captured video set, determine the captured image in the captured video whose frame time corresponds to the video frame time as the target captured image; Generate a key panoramic captured image according to the obtained target captured image set; Package and store the obtained key panoramic captured image sequence and the corresponding key video frame time sequence to obtain a second image package file; Generate second file display information for the second image package file; Display the second file display information on a display terminal for the security system terminal to obtain corresponding key image content; In response to receiving key image content acquisition information for the target object, obtain at least one key panoramic captured image and the corresponding at least one key video frame time for the key image content acquisition information from the stored second image package file; Determine at least one key captured video frame set corresponding to the at least one key video frame time; Send the at least one key panoramic captured image, the at least one key video frame time, and the at least one key captured video frame set to the security system terminal.
3. The method according to claim 1, wherein, After generating the panoramic captured image corresponding to the target video frame time by using the pre-trained generative and adversarial neural network model according to the captured image feature information set and the first image feature weight information set, the method further includes: Input the panoramic captured image into the image feature extraction model to generate panoramic image feature information; Determine the feature difference matrix between the panoramic image feature information and the previous panoramic image feature information; Determine the first difference information between the feature difference matrix and the preset matrix; Input the captured image feature information set and the panoramic image feature information into the image feature weight information generation model to generate an actual image feature weight information set as the actual image feature weight information set; Determine the second difference information between the actual image feature weight information set and the initial image feature weight information set; Input the panoramic captured image and the previous panoramic captured image into a pre-trained image content difference information generation model to generate image content difference information; In response to determining that the first difference information meets the first preset difference condition, the second difference information meets the second preset difference condition, and the image content difference information meets the third preset difference condition, generate verification information indicating that the panoramic captured image passes the verification.
4. The method according to claim 1, wherein, Before the step of, in response to determining that the target video frame time is not the first video frame time in the video frame time sequence, determining the previous video frame time corresponding to the target video frame time, the method further includes: In response to determining that the target video frame time is the first video frame time in the video frame time sequence, determine the captured video frame set corresponding to the target video frame time and the corresponding shooting azimuth information set; Generate a candidate panoramic captured image for the captured video frame set by using panoramic image generation technology; Input the captured image feature information set and the candidate panoramic captured image into the image feature weight information generation model to generate an image feature weight information set as the candidate image feature weight information set; Perform corresponding information splicing on the candidate image feature weight information set, the second captured image feature information set, and the azimuth coding information set to generate a second splicing information set; Input the second splicing information set into the attention mechanism model based on azimuth and feature content to generate a second image feature weight information set for the captured image feature information set; According to the captured image feature information set and the second image feature weight information set, use a pre-trained generative and adversarial neural network model to generate a panoramic captured image corresponding to the target video frame time.
5. The method according to claim 2, wherein The step of screening out a target number of key captured video frames from the captured video frame sequence set as the key captured video frame set includes: Obtain a video usage information set for the captured video set; For each video usage information in the video usage information set, perform the following third generation step: Determine at least one attribute label corresponding to the video usage information; Obtain at least one attribute label information generation model corresponding to the at least one attribute label; For each attribute label information generation model in the at least one attribute label information generation model, perform the following fourth generation step: Input each captured video frame in the captured video frame sequence set into the attribute label information generation model to generate attribute label information, obtaining an attribute label information sequence set, where the attribute label information includes: label name information and label probability information; Screen out the captured video frame with the same corresponding label name information as the target attribute label information and the largest corresponding label probability information from the captured video frame sequence set as the key captured video frame, where the target attribute label information is the attribute label information corresponding to the attribute label information generation model.
6. The method according to claim 1, wherein After the panoramic captured image sequence and the corresponding video frame time sequence are correspondingly packaged and stored to obtain the first image package file, the method further includes: generating first file display information corresponding to the first image package file; displaying the first file display information on a display terminal for the security system terminal to obtain corresponding image content.
7. An information sending device, comprising: a first obtaining unit configured to obtain a set of captured videos for a target object in response to receiving a security information obtaining request for the target object sent by a security system terminal, wherein each captured video in the set of captured videos is captured for each shooting direction; a video preprocessing unit configured to perform video preprocessing on each captured video in the set of captured videos to generate a set of captured video frame sequences; a recombination unit configured to perform video frame recombination on the set of captured video frame sequences according to the video frame time corresponding to each captured video frame to obtain a set sequence of captured video frames, wherein the video frame times corresponding to the respective captured video frames in each set of captured video frames are the same; A generating unit, configured to generate a panoramic captured image sequence for the sequence of captured video frame sets, wherein generating the panoramic captured image sequence for the sequence of captured video frame sets includes: for a target video frame time in the video frame time sequence, performing the following second generating step: in response to determining that the target video frame time is not the first video frame time in the video frame time sequence, determining the previous video frame time corresponding to the target video frame time; determining the panoramic captured image corresponding to the previous video frame time as the previous panoramic captured image; inputting the previous panoramic captured image into a pre-trained image feature extraction model to generate previous panoramic image feature information; determining the captured video frame set corresponding to the target video frame time and the corresponding set of captured orientation information; inputting each captured video frame in the captured video frame set into the image feature extraction model to generate captured image feature information, obtaining a set of captured image feature information; multiplying the previous panoramic image feature information by a preset matrix to generate a multiplication matrix; inputting the set of captured image feature information and the multiplication matrix into a pre-trained image feature weight information generating model to generate a set of image feature weight information as a first initial set of image feature weight information; performing encoding processing on each captured orientation information in the set of captured orientation information to generate a set of orientation encoding information; performing corresponding information splicing on the first initial set of image feature weight information, the set of captured image feature information, and the set of orientation encoding information to generate a first splicing information set; inputting the first splicing information set into an attention mechanism model based on orientation and feature content to generate a first set of image feature weight information for the set of captured image feature information; according to the set of captured image feature information and the first set of image feature weight information, using a pre-trained generative and adversarial neural network model to generate a panoramic captured image corresponding to the target video frame time; in response to determining that the target video frame time is the second video frame time, sorting each panoramic captured image in the obtained panoramic video image set to generate a panoramic video image sequence; in response to determining that the target video frame time is not the second video frame time, determining the next video frame time corresponding to the target video frame time as the target video frame time, and continuing to perform the generating step; A packing and storing unit, configured to correspondingly pack and store the panoramic captured image sequence and the corresponding video frame time sequence to obtain a first image packing file, wherein there is a time correspondence relationship between the video frame time sequence and the sequence of captured video frame sets; A second obtaining unit, configured to, in response to receiving object picture segment obtaining information for the target object, obtain at least one panoramic captured image and the corresponding at least one video frame time for the object picture segment obtaining information from the stored first image packing file, wherein the object picture segment obtaining information includes: object picture obtaining orientation information and object picture obtaining time; A determination unit, configured to determine at least one set of captured video frames corresponding to the at least one video frame time; A sending unit, configured to send the at least one panoramic captured image, the at least one video frame time, and the at least one set of captured video frames to a security system terminal.
8. An electronic device, comprising: One or more processors; A storage device having stored thereon one or more programs, When the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method according to any one of claims 1-6.
9. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Image splicing method, electronic equipment and computer readable storage medium
CN114881863A
Picture director method and device for online class, electronic equipment and storage medium
CN116016978A