Data processing method and device
By acquiring and analyzing the image data collected by multiple devices and dynamically adjusting the video output mode, the flexibility and adaptability of the video output mode of the conference machine or cloud device in multi-person conference scenarios is solved, and a more flexible and stable video display is achieved.
Patent Information
- Application Number
- CN202510467503.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, the video output mode of conference machines or cloud devices has low flexibility and adaptability in multi-person conference scenarios, and it is difficult to dynamically adjust the video output according to the participant's participation.
By acquiring image data collected by multiple devices, the object's participation information for the current event is determined, and the output mode of the image data is dynamically adjusted based on this information, including the recognition of the target object and the classification, segmentation, matching and replacement of the image data, and the video output is optimized.
It improves the flexibility and adaptability of video output, and can adjust the video display in real time according to the participants' participation, reduce video jitter and improve user experience.
Smart Images

Figure CN120455620A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of electronic technology, and in particular to a data processing method and device. Background Art
[0002] In multi-person conference scenarios, the video output mode needs to be pre-set, and the conference machine or cloud will output the collected video data based on this output mode. As can be seen, the flexibility and adaptability of the video output of the conference machine or cloud in related technologies to different scenarios are relatively low. Summary of the Invention
[0003] In view of this, embodiments of the present application provide at least one data processing method.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] In a first aspect, an embodiment of the present application provides a data processing method, the data processing method comprising:
[0006] Acquiring image data collected by at least one device;
[0007] Determining, based on the image data, participation information of an object in the image data with respect to a current event; the participation information is used to characterize a degree of association between the object and the current event;
[0008] determining an output mode of the image data based on the engagement information;
[0009] Based on the output mode, the image data is output.
[0010] In a second aspect, an embodiment of the present application provides a data processing device, the data processing device comprising:
[0011] An acquisition module, configured to acquire image data collected by at least one device;
[0012] A first determining module is configured to determine, based on the image data, participation information of an object in the image data with respect to a current event; the participation information is used to represent a degree of association between the object and the current event;
[0013] a second determining module, configured to determine an output mode of the image data based on the engagement information;
[0014] An output module is configured to output the image data based on the output mode.
[0015] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the technical solutions of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.
[0017] Figure 1 A schematic diagram of the implementation process of a data processing method provided in an embodiment of the present application Figure 1 ;
[0018] Figure 2a Schematic diagram 2 of an implementation flow of a data processing method provided in an embodiment of the present application;
[0019] Figure 2b A scenario diagram of a data processing method provided in an embodiment of the present application Figure 1 ;
[0020] Figure 2c Scenario 2 of a data processing method provided in an embodiment of the present application;
[0021] Figure 2d A scenario diagram of a data processing method provided in an embodiment of the present application Figure 3 ;
[0022] Figure 3 A schematic diagram of the implementation process of a data processing method provided in an embodiment of the present application Figure 3 ;
[0023] Figure 4 A schematic diagram of the implementation process of a data processing method provided in an embodiment of the present application Figure 4 ;
[0024] Figure 5 A schematic diagram of the implementation process of a data processing method provided in an embodiment of the present application Figure 5 ;
[0025] Figure 6 A scenario diagram of a data processing method provided in an embodiment of the present application Figure 4 ;
[0026] Figure 7a A scenario diagram of a data processing method provided in an embodiment of the present application Figure 5 ;
[0027] Figure 7b A scenario diagram of a data processing method provided in an embodiment of the present application Figure 6 ;
[0028] Figure 8 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application;
[0029] Figure 9This is a schematic diagram of a hardware entity of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions of this application are further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0031] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0032] The terms "first / second / third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first / second / third" can be interchanged with a specific order or sequence where permitted so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing this application only and are not intended to limit this application.
[0034] To address the technical issues in the related art, embodiments of the present application provide a data processing method that can be applied to an electronic device, which can be a host device with a processor, or a conference machine in a conference system. In some embodiments, the electronic device can also be an audio-video box (AV box) in the conference system. In other embodiments, the electronic device can also be a cloud device corresponding to the conference software.
[0035] Figure 1 A schematic diagram of the implementation process of a data processing method provided in an embodiment of the present application Figure 1 ,like Figure 1 As shown, the data processing method can be implemented through steps S101 to S104:
[0036] Step S101: Acquire image data collected by at least one device.
[0037] Here, the at least one device may include only the first device, may include at least one first device and at least one second device, or may include only the at least one second device.
[0038] For example, the first device may be a camera corresponding to the host device, and the second device may be a user's terminal device, which may include but is not limited to a smartphone, a tablet computer, a wearable device, a personal computer (PC), a netbook, etc. In some embodiments, the second device may also be an external camera of the terminal device.
[0039] For example, in a conference scenario, the first device may be a camera in a conference system, which collects image data of participants through the camera, and the second device may be a terminal device of the participant, which collects image data of participants through a built-in camera of the terminal device.
[0040] For example, in a conference scenario, each participant's second device is installed with conference software. After each second device enters the current conference session through the conference software, the cloud device collects the participant's image data through the camera of the second device.
[0041] It is understood that in a conference scenario, because the first device and the host device both belong to the conference system, the first connection relationship between the first device and the host device is automatically established after the host device is started based on the historical configuration parameters for the first device. Therefore, when the host device and the first device are started, the first connection relationship between them is automatically established, and image data captured by the first device can be directly transmitted to the host device based on the first connection relationship.
[0042] For the second device, because it belongs to a participant, a second connection relationship between the second device and the host device cannot be automatically established after the second device and the host device are started. During the host device's current operation phase, the second device must first trigger the establishment of a second connection relationship with the host device. Then, after the second device captures the participant's image data, it sends the image data to the host device. In other words, the first device can directly send the captured image data to the host device, while the second device first establishes a communication connection with the host device and then sends the captured image data to the host device. Therefore, the method for obtaining image data for the first device is different from the method for obtaining image data for the second device.
[0043] In an embodiment of the present application, the second device can join the network where the host device is located, thereby triggering the establishment of a communication connection between the two.
[0044] In some embodiments, the second device may also connect to the host device via Bluetooth, thereby triggering the establishment of a communication connection between the two.
[0045] In some embodiments, the second device may also first send the image data to the cloud corresponding to the host device, and then the cloud sends the image data to the host device.
[0046] Step S102: Determine, based on the image data, participation information of an object in the image data with respect to a current event; the participation information is used to represent the degree of association between the object and the current event.
[0047] Here, the current event may be an event that the object in the image data participates in. For example, in a conference scenario, the current event may be a conference; in an interview scenario, the current event may be an interview event.
[0048] It is understood that image data can capture image information of different subjects, and through this information, the subject's body movements, speech, and facial expressions can be determined. This information can be used to determine the subject's participation in the current event.
[0049] For example, when it is determined that the object in the image data is in a speaking state, participation information representing that the object has a high degree of association with the current event can be determined; when it is determined that the expression of the object in the image data is dazed, participation information representing that the object has a low degree of association with the current event can be determined; when it is determined that the amplitude of the body movement of the object in the image data is large, participation information representing that the object has a high degree of association with the current event can be determined.
[0050] Step S103: determining an output mode of the image data based on the participation information.
[0051] Here, the output mode of the image data may refer to a way of outputting the image data.
[0052] It is understood that the output mode of image data is related to the participation of each object in the image data in the current event. For example, when at least two objects among multiple objects have a high degree of participation in the current event, the image data of the at least two objects can be output simultaneously; when only one object among multiple objects has a high degree of participation in the current event, the image data of that object can be output preferentially; when the image data of multiple objects with high participation is output, if it is determined that the participation of an object has decreased, the image data of that object can be stopped from being output, or the image data of that object can be replaced with the image data of another object in the output area of the image data of that object.
[0053] Step S104: output the image data based on the output mode.
[0054] In the embodiment of the present application, the image data may be output according to an output mode corresponding to the object's participation information.
[0055] In an embodiment of the present application, image data captured by at least one device is used to determine the engagement information of an object in the image data with respect to a current event, which indicates the degree of association between the object and the current event. Then, based on the engagement information, an output mode for the image data is determined. Finally, based on the output mode, the image data is output. In this way, the output mode of the image data can be determined in real time based on the engagement information of the object in the image data with respect to the current event, thereby improving the flexibility of video output by a host device or cloud device and its adaptability to different scenarios.
[0056] In some embodiments, the above method may also be implemented through step S105, and correspondingly, the above step S104 may also be implemented through step S1041:
[0057] Step S105 : determining a plurality of target objects and a target image corresponding to each target object based on the image data.
[0058] Here, the target object may be an object in the image data. For example, in a conference scenario, the target object may be a participant. The target image may be an image that can show the face of the target object.
[0059] In an embodiment of the present application, image data from different sources may be classified first to obtain image data sets corresponding to different target objects; and then the target image corresponding to the target object may be determined in the image data sets.
[0060] It is understandable that because the first device and the second device capture images from different perspectives, the target image of the target object may be image data captured by the first device or the second device. Therefore, it is necessary to first classify multiple image data of the same target object together, and then determine the corresponding target image of the target object in the image data set.
[0061] In some embodiments, image data from different sources may be classified based on the anthropomorphic features of the target object in the image data to obtain image data sets corresponding to different objects. For example, the anthropomorphic features may include at least one of the following: skeletal proportions (e.g., arm span, head-to-body ratio, joint spacing), body type (height, weight, body type), clothing color / texture, hairstyle, shoe style, backpack, hat, accessories, etc.
[0062] In some embodiments, the face deflection angle corresponding to the image data may be determined first; and the image data whose face deflection angle satisfies the deflection condition may be determined as the target image corresponding to the target object.
[0063] In some embodiments, when there are multiple target objects in the image data, the image data may be segmented to obtain image sub-data, and then the multiple image sub-data may be classified.
[0064] It is understandable that in a conference scenario, the image acquisition range of the first device is often relatively large and can capture image data of multiple participants, so it is necessary to segment the image data captured by the first device. The second device may also capture image data of multiple participants.
[0065] In some embodiments, the above step S105 can be implemented by steps S1051 and S1052:
[0066] Step S1051 : Compare at least two registered images of the registered objects in the database with the image data collected by the device to obtain at least two image data of each target object; different image data of each target object corresponds to different image collection angles.
[0067] In an embodiment of the present application, at least two registered images corresponding to different registered objects can be obtained from its own database, and then the registered image is compared with each image data, so as to classify the at least two image data according to different objects.
[0068] It is understandable that the image acquisition range of the first device is different from that of the second device, so the perspectives of the images captured by different devices are different. For the same object, it may be captured by both the first device and the second device. At this time, the image data of multiple devices will contain multi-perspective images of the same object. Therefore, in order to classify the multiple image data according to the object, it is necessary to obtain the registered images of multiple pre-registered registered objects from the database and compare them with the image data. Because the perspective of the registered image of the registered object may be different from the perspective of the image data, it is necessary to obtain multi-angle registered images of each registered object and compare them with the image data, so as to improve the accuracy of image comparison.
[0069] In this embodiment of the present application, the registrant can log in to a registration webpage or application through their own terminal device in advance, and then upload their own registration images from multiple perspectives to the host device or cloud device. In other words, when the registrant appears in image data collected by the device, the registrant can be identified from the image data using the pre-registered registration images.
[0070] In an embodiment of the present application, after obtaining the registered image, the human body features of the registered image and the human body features of the image data can be collected first, and then the human body features between the two images can be compared to achieve object classification of the image data.
[0071] Step S1052 : determining a target image corresponding to each target object based on at least two image data of each target object.
[0072] Here, the target image corresponding to the target object may be an image having the best acquisition angle of the target object among the at least two image data, or an image having the highest similarity to the registered image.
[0073] In an embodiment of the present application, for each target object, the face angles of at least two image data of the target object can be determined, and the image data that meets the face deflection angle threshold is determined as the target image.
[0074] In this embodiment of the present application, the number of facial key points in each image data can be determined first, and then first image data whose number of facial key points meets a threshold can be obtained. The facial deflection angle in the first image data can then be determined, and the first image data whose facial deflection angle meets a preset angle can be determined as the target image. For example, the threshold can be 68 or 109.
[0075] It is understandable that when the face is turned sideways, it is impossible to collect all the feature points of the face, so the image data can be filtered by the number of facial key points, and then by calculating the facial deflection angle of the face, the image data that meets the facial deflection angle threshold can be determined.
[0076] Step S1041: outputting at least one target image corresponding to the target object based on the output mode.
[0077] In the embodiment of the present application, when the target object is determined through registered object recognition of a registered object, the target image of the registered object currently captured by the device can be output through the output mode. In other words, even if there are other unregistered objects in the current scene, the image data of the object will not be output.
[0078] In some embodiments, when the target objects are determined based on human body features, all objects can be identified, so the host device or cloud device can output target images corresponding to all target objects based on the output mode.
[0079] In some embodiments, when the target object is determined based on human body features, if the number of target objects in the image data is large, at least one target object will be screened out from multiple target objects based on preset conditions, and a target image corresponding to the at least one target object will be output based on the output mode.
[0080] In some embodiments, as Figure 2aIt is shown that the “determining multiple target objects based on the image data” in the above step S105 can be implemented through steps S201 and S202. When the output mode is the first output mode, the above method can also be implemented through steps S202 to S204:
[0081] Step S201: determining a first speech detection result of an object in the image data based on the image data.
[0082] In the embodiment of the present application, speech detection may be performed on each object in the image data to obtain a first speech detection result corresponding to each object.
[0083] In an embodiment of the present application, human body detection may be performed on image data to obtain a human body image, and then a mouth image of the human body image may be acquired. A first speech detection result corresponding to the object may be determined based on at least two mouth images of the object.
[0084] In some embodiments, when at least one device includes a first device and a second device, and the image acquisition range of the first device is larger than the image acquisition range of the second device, the predicted position information of the speaking object can be determined through the image data collected by the first device, and then the first speech detection result of the object can be determined based on the image data collected by the second device corresponding to the predicted position information.
[0085] It is understandable that the image data collected by the first device is video data, and the video data includes audio data, so the predicted position information of the speaking object can be determined through the audio data.
[0086] Step S202: The object that is speaking indicated by the speech detection result is determined as the target object.
[0087] In the embodiment of the present application, the object that is speaking as indicated by the speech detection result may be determined as the target object, and a target image of the object that is speaking may be output based on the output mode.
[0088] In some embodiments, the target image of the speaking subject may be labeled and then the labeled target image may be output.
[0089] Step S203: determining a second speech detection result of the target object in the output image data.
[0090] In the embodiment of the present application, during the process of continuously outputting the target image of the target object, speech detection is continuously performed on the target object in the output image data to obtain a second speech detection result of the target object in the output image data.
[0091] In the embodiment of the present application, the process of performing speech detection on the target object in the output image data can refer to the implementation of step S201.
[0092] Step S204 : when the second speech detection result indicates that there is a first target object in the output image data that has not spoken, determining a silent time period of the first target object.
[0093] In the embodiment of the present application, when it is determined that there is a first target object that has not spoken in the output image data, the silent time of the first target object can be determined based on multiple target images of the first target object.
[0094] For example, the number of frames of the non-speaking images in the plurality of target images of the first target object may be determined, and the non-speaking time of the first target object may be determined based on the number of frames.
[0095] Step S205, when the duration of not speaking is longer than a predetermined duration, replacing the image data of the first target object with the image data of any of the objects in the output image data;
[0096] The first output mode is a display mode for simultaneously displaying at least two image data. The image data of any object refers to another object different from the object corresponding to the output image data.
[0097] It can be understood that the current output mode is a display mode that displays at least two image data at the same time. Therefore, when it is determined that the object in the output image data has not spoken for a long time, the image data of the object can be replaced with the image data of any object and continue to be output based on the first output mode.
[0098] In the embodiment of the present application, replacing the image data of the first target object with the image data of any object may mean replacing the display position of the image data of the first target object with the image data of another object.
[0099] For example, Figure 2b , four target images are output simultaneously based on the first output mode, the target object in the image data 11 on the far left is the first target object, and when the first target object is silent for longer than a predetermined time, the image data 11 can be replaced by the image data of any object.
[0100] In some embodiments, the “replacing the image data of the first target object with the image data of any of the objects” in step S204 can be implemented through steps S2041 and S2042:
[0101] Step S2041, determining the display position of the image data of the first target object in the output image data; outputting the image data of any of the objects at the display position; or,
[0102] Step S2042: output the image data of any of the objects at an edge position in the output image data.
[0103] For example, Figure 2c , you can Figure 3 Image data 12 is output at the display position of image data 11 in the image data display, thereby replacing image data 11 with image data 12 at the display position of image data 11.
[0104] For example, Figure 2d Image data 12 may be output at the rightmost position in the output image data, and image data 11 may not be output, thereby replacing image data 11 with image data 12 at an edge position in the output image data.
[0105] In some embodiments, as Figure 3 As shown, the above step S104 can be implemented through steps S301 to S303:
[0106] Step S301: determining the position information of each object in the image data in the physical environment.
[0107] In an embodiment of the present application, when at least one device includes a first device, position information of each object in the image data in the physical environment can be determined through image data of the first device.
[0108] For example, in a case where the current scene of the image data is a conference scene, the position information of each object in the physical environment may represent the seat information of each object.
[0109] Step S302: determining an arrangement order or a display order of the image data of each object based on the position information of each object.
[0110] Here, the arrangement order may correspond to a display mode for simultaneously displaying at least two image data (i.e., a first output mode), and the display order may correspond to a display mode for sequentially displaying at least two image data (i.e., a second output mode).
[0111] In an embodiment of the present application, the distance information of each object from the first device can be determined based on the position information of each object, and the arrangement order or display order of the image data of each object can be determined according to the size relationship between the distance information corresponding to each object.
[0112] For example, Figure 2b As shown, the distance information corresponding to the objects in the image data from left to right can be from small to large or from large to small.
[0113] In the embodiment of the present application, the image data corresponding to the object with the smallest distance information can be determined as the image data to be displayed first, and the image data corresponding to the object with the largest distance information can be determined as the image data to be displayed last. In other words, the image data with the smaller distance information has the higher priority to be output.
[0114] Step S303: outputting the image data of the various objects simultaneously according to the arrangement order, or outputting the image data of the various objects sequentially according to the display order.
[0115] In the embodiment of the present application, the image data of each object can be output in an arrangement order based on the first output mode, or the image data of each object can be output in a display order based on the second output mode.
[0116] In some embodiments, the above method may also be implemented through step S11:
[0117] Step S11 : when the relative positions of the objects in the physical environment change, continue to output the image data of the objects simultaneously according to the arrangement sequence.
[0118] In an embodiment of the present application, the background area of the image data can be used to determine whether the relative positions of the various objects in the physical environment have changed. If the relative positions of the various objects in the physical environment have changed, the above arrangement order will not be changed, and the image data of the various objects will continue to be output simultaneously according to the arrangement order.
[0119] In the embodiment of the present application, when the relative positions of the objects in the physical environment change, the image data of the objects continue to be output simultaneously in the arrangement order, thereby reducing the jitter caused by the change in the output position of the objects.
[0120] In some embodiments, as Figure 4 As shown, the above method can also be implemented through steps S401 to S405. Correspondingly, the above step S104 can be implemented through step S406:
[0121] Step S401 : performing image matching on a target image corresponding to the target object in a first database including registered objects and / or a second database including unregistered objects to obtain a matching result.
[0122] In the embodiment of the present application, a first database is pre-stored with multiple registered images of registered objects, and a second database is pre-stored with multiple images of unregistered objects. Then, the target image of the target object is matched with the first database and the second database.
[0123] It is understood that a registered user may pre-register on a host device or cloud device, thereby obtaining multiple registered images of registered objects. Regarding the second database, when the host device or cloud device determines the target image, images other than the registered image may be identified as images of unregistered objects and then stored in the second database.
[0124] Step S402 : when the matching result represents the target image and the similarity between the target image and the image in the first database is greater than or equal to a predetermined threshold, determining the attribute information of the registered object corresponding to the first database as the identification information of the target image.
[0125] Exemplarily, the attribute information of the registration object may include the name and title of the registration object.
[0126] In this embodiment of the present application, a first database stores correspondences between different registered images and attribute information of different registered objects. A registered image whose similarity to a target image is greater than or equal to a predetermined threshold is identified. Based on this registered image, the attribute information of the corresponding registered object is then determined from the multiple correspondences. Finally, the attribute information of the registered object is determined as identification information for the target image.
[0127] Step S403 : when the matching result represents the target image and the similarity between the target image and the image in the second database is greater than or equal to a predetermined threshold, determining the unregistered object information as identification information of the target image.
[0128] Here, the unregistered object information may indicate that the object in the image data is not registered in the host device or the cloud device. For example, the unregistered object information may be “guest” or “guest”.
[0129] In an embodiment of the present application, when there is an image in the second database whose similarity with the target image is greater than or equal to a predetermined threshold, the object representing the target image is not registered in the host device or the cloud device, so the unregistered object information can be determined as the identification information of the target image.
[0130] Step S404 : when the matching result represents the target image and the similarity between the target image and each image in the first database and the second database is less than a predetermined threshold, determining the unregistered object information as identification information of the target image.
[0131] In an embodiment of the present application, if there is no image in the first database and the second database whose similarity with the target image is greater than or equal to a predetermined threshold, that is, the similarity between the target image and each image in the first database and the second database is less than the predetermined threshold, it may also mean that the object of the target image has not been registered in the host device or the cloud device, and therefore the unregistered object information can be determined as the identification information of the target image.
[0132] In some embodiments, the target image whose similarity with each image in the first database and the second database is less than a predetermined threshold can be stored in the second database to complete the update of the second database.
[0133] Step S405 : Mark the target image with the identification information to obtain a processed target image.
[0134] Step S406: output the processed target image based on the output mode.
[0135] In the embodiment of the present application, corresponding identification information can be marked on each target image to be output, and then the processed target image can be output in the current output mode.
[0136] In some embodiments, the above step S102 may be implemented by steps S21 and S22:
[0137] Step S21 : determining the behavior status information of the object based on the image data; the behavior status information includes at least one of the following: body movements, speaking conditions, and facial expressions.
[0138] In an embodiment of the present application, at least one of the following detections can be performed on the object in the image data: body movement detection, speech detection, and expression detection, and at least one of the detected body movement, speech situation, and expression state of the object is determined as the behavioral state information of the object.
[0139] Step S22: When at least one of the subject's body movements, speech, and facial expressions satisfies a participation condition, determining the subject's participation information as first participation information.
[0140] Here, the first participation information may indicate that the object's relevance to the current event is greater than a preset value. That is, when the object's participation information is the first participation information, the object's relevance to the current event is high, and the object's participation in the current event is also high.
[0141] In the embodiment of the present application, the participation conditions include a first participation condition corresponding to body movements, a second participation condition corresponding to speaking situations, and a third participation condition corresponding to facial expression states.
[0142] Regarding the first participation condition: when it is determined that the object's body movement is a hand-raising state, it can be determined that the object's body movement meets the first participation condition; when it is determined that the object's body movement is a communication gesture, it can be determined that the object's body movement meets the first participation condition.
[0143] Regarding the second participation condition: when the subject's speaking situation indicates that the subject is speaking, it can be determined that the subject's speaking situation meets the second participation condition.
[0144] Regarding the third participation condition: when it is determined that the expression state of the object is a listening state, it can be determined that the expression state of the object meets the third participation condition.
[0145] In some embodiments, the above step S103 can be implemented by steps S31 and S32:
[0146] Step S31: When the number of objects corresponding to the first participation information is at least two, determine that the output mode is the first output mode.
[0147] Here, the first output mode is a display mode for simultaneously displaying at least two image data.
[0148] It is understandable that when the engagement information of multiple objects in the image data is the first engagement information, it can be determined that the multiple objects are highly associated with the current event. In this case, the output mode of the host device or the cloud device can be determined to be the first output mode, that is, the image data of multiple objects whose engagement information is the first engagement information are displayed simultaneously, for example Figure 2b .
[0149] Step S32: When the number of objects corresponding to the first participation information is one, determine that the output mode is the second output mode.
[0150] Here, the second output mode is a display mode for sequentially displaying at least two image data.
[0151] It can be understood that because the number of objects corresponding to the current first engagement information is one, in order to focus on the object, the output mode of the host device or the cloud device can be determined to be a display mode that displays at least two image data in sequence, and the image data of the object whose engagement information is the first engagement information is output preferentially.
[0152] In some embodiments, determining the output mode of the host device or cloud device is ongoing. When the output mode of the host device or cloud device is the first output mode, and it is determined that the body movements, speech, and facial expressions of the subject in the output image data do not meet the participation conditions, the host device or cloud device may switch to the second output mode. Similarly, when the output mode of the host device or cloud device is the second output mode, and there are multiple subjects in the image data whose participation information is the first participation information, the host device or cloud device may switch to the first output mode.
[0153] The following describes the application of the data processing method provided in the embodiment of the present application in a practical scenario:
[0154] Face recognition technology on personal computers (PCs) is very mature, mainly because the face and the PC have a good angle and are very close, which is suitable for face recognition scenarios.
[0155] In conference room video scenarios, it is quite difficult to identify people in the conference room. The face of each conference room participant does not have a fixed angle relationship with the camera, making it difficult to obtain the best facial angle. Moreover, if the person is far away from the camera, the facial details are limited, resulting in a decrease in accuracy.
[0156] In addition, in the video gallery scene (a layout where multiple video screens are arranged and displayed), it often happens that even if the displayed people do not change, their positions in the video change randomly, causing video jitter and affecting the final experience. With the help of character recognition technology, the effect of anti-video jitter can be achieved.
[0157] In order to solve the technical problems in the related art, the embodiment of the present application provides a data processing method, the execution subject of the method can be a conference machine or an audio-video control box (Audio-Video Box, AV Box), such as Figure 5 As shown, the data processing method can be implemented through steps S1 to S8:
[0158] Step S1: Acquire image data collected by multiple devices.
[0159] For example, Figure 6 As shown, in the current conference scenario, the number of participants is 4, and image 2111 is captured by the terminal device 21 corresponding to the participant, and image 2111 is sent to the conference machine 31; the conference machine camera 311 can capture image 3111 and send image 3111 to the conference machine 31.
[0160] Step S2: human body detection.
[0161] Step S3, re-identification.
[0162] In the embodiment of the present application, when a user exists in the image data, re-identification is performed, that is, image data of the same user are classified together based on pre-stored user information.
[0163] Step S4: Obtain updated user features.
[0164] In the embodiment of the present application, the conference machine can obtain updated user features from the user information in the database, wherein the user features include user features and guest features.
[0165] For example, Figure 7a and Figure 7b As shown, users can upload their own images, names, and positions on the registration page 41 through their own terminal devices. Among them, their own images can be user images under different perspectives, such as Figure 6 As shown, it includes user images 42, 43, and 44 from three perspectives. The conference machine stores the uploaded images in database 45. During use, when at least one device 46 (including the conference machine's camera and the participant's own device) captures a participant's image data, it can upload the image data to the conference machine. The conference machine then uses database 45 to perform image matching on the captured image data and obtain output 46. Simultaneously, the conference machine can also use the captured image data to update the image data in database 45 through data management.
[0166] Step S5: determining the human body features in the image as user features based on the updated user features.
[0167] Step S6: determining the human body features in the image as guest features based on the updated user features.
[0168] Step S7: Feedback the user features and guest features in the image to the database.
[0169] Step S8: Obtain the re-identification result.
[0170] Figure 8 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application is shown in FIG. Figure 8 As shown, the data processing device 800 includes an obtaining module 801, a first determining module 802, a second determining module 803 and an output module 804; wherein,
[0171] An acquisition module 801 is configured to acquire image data collected by at least one device;
[0172] A first determining module 802 is configured to determine, based on the image data, participation information of an object in the image data with respect to a current event; the participation information is used to represent a degree of association between the object and the current event;
[0173] A second determining module 803 is configured to determine an output mode of the image data based on the engagement information;
[0174] The output module 804 is configured to output the image data based on the output mode.
[0175] In some embodiments, the above-mentioned data processing device also includes a third determination unit, which is used to determine multiple target objects and target images corresponding to each target object based on the image data; the above-mentioned output module 804 is also used to output at least one target image corresponding to the target object based on the output mode.
[0176] In some embodiments, the third determination module is further used to determine a first speech detection result of an object in the image data based on the image data; the object that is speaking indicated by the speech detection result is determined as the target object; when the output mode is the first output mode, the above-mentioned data processing device also includes a fourth determination module, which is used to determine a second speech detection result of the target object in the output image data; when the second speech detection result indicates that there is a first target object that has not spoken in the target object in the output image data, determine the duration of silence of the first target object; when the duration of silence is greater than a predetermined duration, replace the image data of the first target object with the image data of any of the objects in the output image data; wherein, the first output mode is a display mode that displays at least two image data at the same time.
[0177] In some embodiments, the fourth determination module is also used to determine the display position of the image data of the first target object in the output image data; output the image data of any of the objects at the display position; or output the image data of any of the objects at the edge position in the output image data.
[0178] In some embodiments, the output module 804 is further used to determine the position information of each object in the image data in the physical environment; based on the position information of each object, determine the arrangement order or display order of the image data of each object; output the image data of each object simultaneously according to the arrangement order, or output the image data of each object in sequence according to the display order.
[0179] In some embodiments, the output module 804 is further configured to continue to output the image data of the objects simultaneously in the arrangement order when the relative positions of the objects in the physical environment change.
[0180] In some embodiments, the data processing device further includes a matching module, a fifth determination module, and an identification module; the matching module is configured to perform image matching on the target image corresponding to the target object in a first database including registered objects and / or a second database including unregistered objects to obtain a matching result; the fifth determination module is configured to determine the attribute information of the registered object corresponding to the first database as the identification information of the target image when the matching result represents that the target image has a similarity greater than or equal to a predetermined threshold with the image in the first database; determine the unregistered object information as the identification information of the target image when the matching result represents that the target image has a similarity greater than or equal to a predetermined threshold with the image in the second database; determine the unregistered object information as the identification information of the target image when the matching result represents that the target image has a similarity less than a predetermined threshold with each image in the first database and the second database; the identification module is configured to identify the identification information on the target image to obtain a processed target image; the output module 804 is further configured to output the processed target image based on the output mode.
[0181] In some embodiments, the first determination module 802 is further used to determine the behavioral status information of the object based on the image data; the behavioral status information includes at least one of the following: body movements, speaking conditions, and facial expressions; when at least one of the body movements, speaking conditions, and facial expressions of the object meets the participation conditions, the participation information of the object is determined to be the first participation information.
[0182] In some embodiments, the second determination module 803 is further used to determine that the output mode is the first output mode when the number of objects corresponding to the first engagement information is at least two; and to determine that the output mode is the second output mode when the number of objects corresponding to the first engagement information is one; wherein the first output mode is a display mode that displays at least two image data simultaneously; and the second output mode is a display mode that displays at least two image data sequentially.
[0183] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to perform the methods described in the above method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0184] It should be noted that, in the embodiment of the present application, if the above-mentioned data processing method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the relevant technology can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a magnetic disk or an optical disk. In this way, the embodiment of the present application is not limited to any specific hardware, software or firmware, or any combination of hardware, software and firmware.
[0185] An embodiment of the present application provides a computer device including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.
[0186] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above method. The computer-readable storage medium may be transient or non-transient.
[0187] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code runs in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.
[0188] The present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented. The computer program product can be implemented in hardware, software, or a combination thereof. In some embodiments, the computer program product is embodied as a computer storage medium. In other embodiments, the computer program product is embodied as a software product, such as a software development kit (SDK).
[0189] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between the various embodiments, and their similarities or similarities can be referenced to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the description of the method embodiments of this application for understanding.
[0190] Figure 9 This is a hardware entity diagram of an electronic device in an embodiment of the present application, such as Figure 9 As shown, the hardware entity of the electronic device 900 includes: a processor 901, a communication interface 902 and a memory 903, wherein:
[0191] The processor 901 generally controls the overall operation of the electronic device 900 , and the overall operation may be to implement the data processing method provided in the embodiment of the present application.
[0192] The communication interface 902 enables the electronic device 900 to communicate with other terminals or servers through a network.
[0193] The memory 903 is configured to store instructions and applications executable by the processor 901. It can also cache data to be processed or processed by the processor 901 and various modules in the electronic device 900 (for example, image data, audio data, voice communication data, and video communication data). This can be implemented using flash memory (FLASH) or random access memory (RAM). Data can be transmitted between the processor 901, the communication interface 902, and the memory 903 via a bus 904.
[0194] An embodiment of the present application provides a computer storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps of the data processing method of any of the above embodiments.
[0195] It should be noted that the description of the above storage medium and device embodiments is similar to the description of the above method embodiments and has similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium and device embodiments of this application, please refer to the description of the method embodiments of this application for understanding.
[0196] The processor may be at least one of an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a programmable logic device (PLD), a field programmable gate array (FPGA), a central processing unit (CPU), a controller, a microcontroller, and a microprocessor. It is understood that the electronic device that implements the functions of the processor may also be other electronic devices, which are not specifically limited in the embodiments of the present application.
[0197] The above-mentioned computer storage medium / memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory (Flash Memory), a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various terminals including one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.
[0198] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.
[0199] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0200] The above are only implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the protection scope of the present application.
Claims
1. A data processing method, comprising: Acquiring image data collected by at least one device; Determining, based on the image data, participation information of an object in the image data with respect to a current event; The participation information is used to represent the degree of association between the object and the current event; determining an output mode of the image data based on the engagement information; Based on the output mode, the image data is output.
2. The method according to claim 1, further comprising: determining a plurality of target objects and a target image corresponding to each target object based on the image data; Outputting the image data based on the output mode includes: Based on the output mode, at least one target image corresponding to the target object is output.
3. The method according to claim 2, wherein determining a plurality of target objects based on the image data comprises: determining, based on the image data, a first speech detection result of an object in the image data; The object that is speaking indicated by the speech detection result is determined as the target object; When the output mode is the first output mode, the method further includes: determining a second speech detection result of the target object in the output image data; If the second speech detection result indicates that there is a first target object in the output image data that has not spoken, determining a silence time of the first target object; When the silent period is longer than a predetermined period, replacing the image data of the first target object with the image data of any of the objects in the output image data; The first output mode is a display mode for simultaneously displaying at least two image data.
4. The method according to claim 3, wherein replacing the image data of the first target object with the image data of any of the objects comprises: determining a display position of the image data of the first target object in the output image data; outputting the image data of any of the objects at the display position; or, The image data of any of the objects is output at an edge position in the output image data.
5. The method according to claim 1 or 2, wherein outputting the image data based on the output mode comprises: Determining position information of each object in the image data in the physical environment; determining an arrangement order or a display order of the image data of the respective objects based on the position information of the respective objects; The image data of the various objects are output simultaneously according to the arrangement order, or the image data of the various objects are output sequentially according to the display order.
6. The method according to claim 5, further comprising: When the relative positions of the objects in the physical environment change, the image data of the objects continue to be output simultaneously according to the arrangement sequence.
7. The method according to claim 2, further comprising: Performing image matching on a target image corresponding to the target object in a first database including registered objects and / or a second database including unregistered objects to obtain a matching result; If the matching result represents the target image and the similarity between the target image and the image in the first database is greater than or equal to a predetermined threshold, determining the attribute information of the registered object corresponding to the first database as the identification information of the target image; If the matching result represents the target image and the similarity between the target image and the image in the second database is greater than or equal to a predetermined threshold, determining the unregistered object information as the identification information of the target image; If the matching result represents the target image and the similarity between the target image and each image in the first database and the second database is less than a predetermined threshold, determining the unregistered object information as the identification information of the target image; Marking the identification information on the target image to obtain a processed target image; Outputting at least one target image corresponding to the target object based on the output mode includes: Based on the output mode, the processed target image is output.
8. The method according to any one of claims 1 to 4, wherein determining, based on the image data, information about the degree of participation of an object in the image data in a current event comprises: determining behavioral status information of the object based on the image data; The behavior status information includes at least one of the following: body movements, speech status, and facial expression status; When at least one of the subject's body movements, speech situations, and facial expressions satisfies a participation condition, the subject's participation information is determined to be first participation information.
9. The method according to claim 8, wherein determining an output mode of the image data based on the engagement information comprises: When the number of objects corresponding to the first engagement information is at least two, determining that the output mode is the first output mode; When the number of objects corresponding to the first participation information is one, determining the output mode to be the second output mode; The first output mode is a display mode for displaying at least two image data simultaneously; and the second output mode is a display mode for displaying at least two image data sequentially.
10. A data processing device comprising: An acquisition module, configured to acquire image data collected by at least one device; A first determining module is configured to determine, based on the image data, participation information of an object in the image data with respect to a current event; The participation information is used to represent the degree of association between the object and the current event; a second determining module, configured to determine an output mode of the image data based on the engagement information; An output module is configured to output the image data based on the output mode.