Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing at least one program, and a method for managing chat conversation retention models. The method may include causing display of an interface that enables a user to select one of multiple retention models for association with an electronic chat conversation, and receiving, via the selector interface, a selection of a particular retention model. The retention model specifies an amount of time that each individual message in the electronic chat conversation is accessible upon being read by a receiving user. The method further includes storing a newly received message as part of the chat conversation, where the storing includes configuring a retention duration attribute for the message in accordance with the amount of time specified by the retention model. The method further includes erasing the message in accordance with the retention duration attribute.
H04L 51/42 - Mailbox-related aspects, e.g. synchronisation of mailboxes
H04L 51/04 - Real-time or near real-time messaging, e.g. instant messaging [IM]
H04L 51/216 - Handling conversation history, e.g. grouping of messages in sessions or threads
H04L 51/52 - User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail for supporting social networking services
Systems and methods are provided for generating dynamic profile backgrounds in a social networking application. A system receives contextual data from a client device including location data, weather data, time data, and event data. The system generates a prompt by inserting contextual parameters into a parameterized prompt template and provides the prompt to a generative artificial intelligence model to generate a background image. The generated background image can be combined with visual effect overlays selected based on weather conditions or events. The system displays the background image with a graphical avatar in a user interface and updates the background on a rolling basis in response to changes in contextual data while adhering to update frequency limitations. The system can store pre-generated assets for major locations to reduce processing requirements and provides fallback behavior when location services are disabled.
A closed loop control system actively regulates the battery current paths of physically separated circuits so that the current is approximately the same for each of the circuits regardless of the various system loads. The closed loop control system modulates the current paths by either modulating a high side transistor used to independently limit each battery's current path or by modulating a DC/DC converter's output voltage to independently boost each battery's current path. The closed loop control system is also designed to handle undervoltage lockout (UVLO) situations when one of the batteries is nearing empty to tilt the power balance in the chance that there is an existing battery charge mismatch to support system load bursts and to turn off the circuit when the system current draw is exceptionally low. A tilting circuit also identifies and discharges the battery with the higher charge until the charge states are substantially equal.
Collaborative sessions in which access to a collaborative object and added virtual content is selectively provided to participants/users. In one example of the collaborative session, a participant crops media content by use of a hand gesture to produce an image segment that can be associated to the collaborative object. The hand gesture resembles a pair of scissors and the camera and processor of the client device track a path of the hand gesture to identify an object within a displayed image to create virtual content of the identified object. The virtual content created by the hand gesture is then associated to the collaborative object.
Examples relate to systems and methods for providing a machine learning model on a mobile device. The systems and methods store a base machine learning (ML) model on a user system, the base ML model trained to perform a first task in an individual domain. The systems and methods receive input that selects a second task in the individual domain and access parameter update information associated with the second task. The systems and methods update the base ML model based on the parameter update information associated with the second task and generate an output corresponding to the second task by processing an input by the updated base ML model.
Optical systems and lens assemblies suitable for use with, for example, AR applications on portable electronic devices, including wearable devices such as electronic eyewear. The lens assembly supports a single lens or multiple lenses. The lens assembly includes a lens and a flange that defines a recess for receiving an adhesive. When the lens is placed on the flange, the adhesive secures the lens to the flange. In some implementations, the lens includes a band or strip of paint or other indicia applied along the perimeter of the lens to inhibit visibility of any excess adhesive. A lens assembly that supports both a front and rear lens is particularly useful for presenting AR experiences.
A carry case for an electronics-enabled eyewear device has a case body defining a storage chamber between a pair of opposing main walls configured to bear against a front and a rear of the eyewear device when received in the storage chamber. A power source is housed by the case body. Two or more charging surfaces are electrically connected to the power source and located in the storage chamber for contact engagement with complementary contact formations on the eyewear device to enable charging of an onboard battery of the eyewear device. Each of the opposing main walls has mounted thereon at least one charging surface that extends along the respective main wall for a majority of a height dimension of the main wall. In some embodiments, each main wall carries a pair of charging pads spaced apart along a length dimension of the storage chamber, the charging pads being arranged such that the eyewear device is chargeable in any one of four different orientations in the storage chamber.
H02J 50/10 - Circuit arrangements or systems for wireless supply or distribution of electric power using inductive coupling
H02J 50/70 - Circuit arrangements or systems for wireless supply or distribution of electric power involving the reduction of electric, magnetic or electromagnetic leakage fields
8.
Lossy video encoding for battery-constrained devices
Systems, methods, and computer readable media for lossy video encoding for battery-constrained devices where the methods performed on a system include accessing a gaze location of a user viewing a first frame of a video on a wearable device, determining a first portion of a second frame based on the gaze location, compressing, in accordance with a first lossy compression standard, the first portion of the second frame to generate a first compressed portion of the second frame, compressing, in accordance with a second lossy compression standard a second portion of the second frame, to generate a second compressed portion of the second frame, the second portion comprising a portion of the second frame not included in the first portion, and causing the first compressed portion and the second compressed portion to be sent wirelessly to the wearable device.
H04N 19/172 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being an image region, e.g. an object the region being a picture, frame or field
H04N 19/20 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using video object coding
H04N 19/42 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by implementation details or hardware specially adapted for video compression or decompression, e.g. dedicated software implementation
H04N 19/463 - Embedding additional information in the video signal during the compression process by compressing encoding parameters before transmission
Methods and systems are disclosed for providing shortcuts to sharing private collections of content items. The methods and systems receive, by a device associated with a first user, a request to share a first content item with one or more recipients and in response, presenting a plurality of shortcuts, each of the plurality of shortcuts associated with different groups of recipients. The methods and systems receive input that selects a first shortcut of the plurality of shortcuts, the first shortcut being associated with a first group of recipients and, in response, present an option to add the first content item to a collection of content items and share the collection of content items with the first group of recipients.
In various embodiments, an optical system to present an image to an eye of a user is disclosed. The system comprises a waveguide configured to output collimated light towards an optically powered element comprising at least one holographic component to generate optical power. The optically powered element is configured to receive the output collimated light from the waveguide and direct the received light towards the eye of the user and impart an angular offset on the directed light such that the directed light forms a virtual image plane.
A message composition system to generate and distribute a plurality of messages to individual recipients based on a single message request, wherein each message among the plurality of messages is addressed and sent to a distinct recipient. According to certain embodiments, the message composition system is configured to perform operations that include, receiving a request to generate a message at a client device, causing display of a composition interface in response to the request to generate the message, wherein the composition interface includes a presentation of a menu that includes a list of user contacts, receiving an identification of a plurality of user contacts from among the list of user contacts, and generating a set of messages in response to the identification of the plurality of user contacts, wherein the set of messages are each individually addressed to the users among the plurality of user contacts.
H04L 51/52 - User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail for supporting social networking services
H04L 51/04 - Real-time or near real-time messaging, e.g. instant messaging [IM]
12.
AUTOMATIC ELECTRON BEAM CALIBRATION ON PERIODIC NANOSTRUCTURES
Systems, methods, and media for calibrating a scanning electron microscope (SEM). The system includes a processor and a memory. The memory stores instructions that, when executed by the processor, configure the system to perform operations. An electron microscope image of a periodic structure is generated by the SEM. A Fourier transform of the electron microscope image is computed to generate a spectrum. Reciprocal lattice vectors are computed based on a known periodicity of the periodic structure. A pixel mask is generated based on the reciprocal lattice vectors and applied to filter the spectrum. A quality metric is generated based on an aggregate magnitude of the filtered spectrum and a magnitude of a zero-frequency component of the filtered spectrum. A pixel scaling parameter, focus parameter, and/or stigmation parameter of the SEM are determined based on the quality metric.
Systems and methods are disclosed for pruning artificial deep neural networks. The systems and methods receive input comprising a target latency associated with executing a machine learning model on a specified device. The systems and methods modify a set of parameters of the machine learning model to reduce a latency associated with the machine learning model based on the target latency. The systems and methods apply the machine learning model with the modified set of parameters to a data set to generate an output.
Systems and methods are provided for generating dynamic profile backgrounds in a social networking application. A system receives contextual data from a client device including location data, weather data, time data, and event data. The system generates a prompt by inserting contextual parameters into a parameterized prompt template and provides the prompt to a generative artificial intelligence model to generate a background image. The generated background image may be combined with visual effect overlays selected based on weather conditions or events. The system displays the background image with a graphical avatar in a user interface and updates the background on a rolling basis in response to changes in contextual data while adhering to update frequency limitations. The system may store pre-generated assets for major locations to reduce processing requirements and provides fallback behavior when location services are disabled.
G06T 11/60 - Editing figures and textCombining figures or text
G06F 3/04845 - Interaction techniques based on graphical user interfaces [GUI] for the control of specific functions or operations, e.g. selecting or manipulating an object, an image or a displayed text element, setting a parameter value or selecting a range for image manipulation, e.g. dragging, rotation, expansion or change of colour
G06T 5/50 - Image enhancement or restoration using two or more images, e.g. averaging or subtraction
Systems and methods provide for editing of spherical video data. In one example, a computing device can receive a spherical video (or a video associated with an angular field of view greater than an angular field of view associated with a display screen of the computing device), such as by a built-in spherical video capturing system or acquiring the video data from another device. The computing device can display the spherical video data. While the spherical video data is displayed, the computing device can track the movement of an object (e.g., the computing device, a user, a real or virtual object represented in the spherical video data, etc.) to change the position of the viewport into the spherical video. The computing device can generate a new video from the new positions of the viewport.
G11B 27/031 - Electronic editing of digitised analogue information signals, e.g. audio or video signals
G06F 3/01 - Input arrangements or combined input and output arrangements for interaction between user and computer
G06F 3/0346 - Pointing devices displaced or positioned by the userAccessories therefor with detection of the device orientation or free movement in a 3D space, e.g. 3D mice, 6-DOF [six degrees of freedom] pointers using gyroscopes, accelerometers or tilt-sensors
G06T 3/16 - Spatio-temporal transformations, e.g. video cubism
G06T 3/4038 - Image mosaicing, e.g. composing plane images from plane sub-images
H04N 5/77 - Interface circuits between an apparatus for recording and another apparatus between a recording apparatus and a television camera
H04N 5/92 - Transformation of the television signal for recording, e.g. modulation, frequency changingInverse transformation for playback
H04N 9/804 - Transformation of the television signal for recording, e.g. modulation, frequency changingInverse transformation for playback involving pulse code modulation of the colour picture signal components
H04N 9/82 - Transformation of the television signal for recording, e.g. modulation, frequency changingInverse transformation for playback the individual colour picture signal components being recorded simultaneously only
Described is a system for identifying content augmentations based on an interaction function initiated by a user by determining an initiation of an interaction function from a first user of an interaction system, processing data associated with the interaction function using a first machine learning model to generate a feature vector, and identifying at least one recommended content augmentation based on a comparison of the feature vector for the interaction function to a feature vector for the at least one recommended content augmentation. The system then displays the at least one recommended content augmentation to the first user with a corresponding selectable user interface element for individual recommended content augmentations.
G06V 10/74 - Image or video pattern matchingProximity measures in feature spaces
G06V 10/77 - Processing image or video features in feature spacesArrangements for image or video recognition or understanding using pattern recognition or machine learning using data integration or data reduction, e.g. principal component analysis [PCA] or independent component analysis [ICA] or self-organising maps [SOM]Blind source separation
A composite image generation system segments a face from at least a first image and a second image and generates a composite image comprising the segmented face of at least the first image and the second image. The composite image generation system further generates a text prompt comprising randomized characteristics for a new image an selects a pose from a plurality of predefined poses. Using a machine learning model, the composite image generation system generates a new image based on the text prompt, the composite image, and the selected pose. The new image can be displayed on a computing device, added or posted to a media collection, other otherwise displayed or shared to one or more users.
Examples relate to systems and methods for modifying a video during a video call. The system receives a request to initiate a video call to a recipient. The system, in response to receiving the request, displays, while attempting to establish a connection with the recipient, a video preview interface that includes a plurality of digital effect indicators available for application during the video call. The system enables selection and application of an individual digital effect indicator from the plurality of digital effect indicators to modify the video preview using a corresponding digital effect.
G06F 3/0482 - Interaction with lists of selectable items, e.g. menus
G06F 3/0484 - Interaction techniques based on graphical user interfaces [GUI] for the control of specific functions or operations, e.g. selecting or manipulating an object, an image or a displayed text element, setting a parameter value or selecting a range
H04L 65/1069 - Session establishment or de-establishment
H04L 65/1089 - In-session procedures by adding mediaIn-session procedures by removing media
19.
CHARGING AND DATA ACCESS PORT FOR HEAD-WEARABLE APPARATUS
In some examples, a head-wearable apparatus for viewing augmented reality (AR) or virtual reality (VR) content is provided. An example the apparatus comprises a frame, an optical assembly including an image display in which the AR or VR content may be viewed by a user, and a user input device operable by the user to navigate through content viewed in the image display, or to invoke a function of the head-wearable apparatus. The user input device includes a body manually engageable by the user to perform a content navigation or function invocation operation and is configured to present at least one contact for accepting a connection to an external charging source, or a connection to an external device.
Examples described herein relate to hand-based light estimation for extended reality (XR). An image sensor of an XR device is used to obtain an image of a hand in a real-world environment. At least part of the image is processed to detect a pose of the hand. One of a plurality of machine learning models is selected based on the detected pose. At least part of the image is processed via the machine learning model to obtain estimated illumination parameter values associated with the hand. The estimated illumination parameter values are used to render virtual content to be presented by the XR device.
The described system facilitates personalized avatar interactions by dynamically generating and displaying modified garments. The system determines the initiation of an interaction function by a first user with a second user within an interaction platform. It accesses avatar data for the first user, including visual attributes and a garment associated with the first avatar, as well as avatar data for the second user, including visual attributes and a garment associated with the second avatar. An image is generated featuring the first avatar wearing its garment alongside the second avatar wearing its garment. This image is applied to the first avatar's garment, creating a modified version that visually incorporates the second avatar. The system then displays the first avatar wearing the modified garment alongside the second avatar wearing its original garment, enabling enhanced user engagement through visually personalized and contextual representations of user interactions.
A63F 13/213 - Input arrangements for video game devices characterised by their sensors, purposes or types comprising photodetecting means, e.g. cameras, photodiodes or infrared cells
A63F 13/53 - Controlling the output signals based on the game progress involving additional visual information provided to the game scene, e.g. by overlay to simulate a head-up display [HUD] or displaying a laser sight in a shooting game
A63F 13/67 - Generating or modifying game content before or while executing the game program, e.g. authoring tools specially adapted for game development or game-integrated level editor adaptively or by learning from player actions, e.g. skill level adjustment or by storing successful combat sequences for re-use
A63F 13/795 - Game security or game management aspects involving player-related data, e.g. identities, accounts, preferences or play histories for finding other playersGame security or game management aspects involving player-related data, e.g. identities, accounts, preferences or play histories for building a teamGame security or game management aspects involving player-related data, e.g. identities, accounts, preferences or play histories for providing a buddy list
A63F 13/87 - Communicating with other players during game play, e.g. by e-mail or chat
Systems and methods for projecting each of a chronology of images as a sequence of images using a shifting element as part of a near-eye display system are provided for use in virtual reality, augmented reality, or mixed reality systems. In some example embodiments, a chronology of images is received by a peripheral sequencing system. The system divides each image into image portions and generates sequences of image portions to recreate the images based on arrangement data. The system then causes a high-speed display of each sequence of images such that they appear simultaneous to a viewer. In some embodiments, the projection is transmitted to a shifting optical element such as a rotating micromirror that propagates a display to a user. In some embodiments, the system further detects and corrects for image and environmental distortions.
An eXtended Reality (XR) content creation environment is provided. A machine learning model is trained by receiving a set of three-dimensional (3D) assets. The training process includes an alignment filter component that generates aligned 3D assets by orienting each 3D asset to face a consistent predefined direction in global-space. A renderer generates multiview global-space normal maps that maintain consistent orientation independently of camera position for the aligned 3D assets. The renderer also generates corresponding renders that serve as ground truth data. The machine learning model is trained using the global-space normal maps and prompts as conditioning input while using the renders as ground truth, with loss backpropagation applied between generated output textures and the ground truth renders.
Eyewear that includes a frame supporting an optical element. The frame has a first side and a second side. The eyewear also includes a temple adjacent the first side of the frame. The temple includes a first portion adjacent the frame and a second portion releasably connected to the first portion. The eyewear also includes an electrical connector embedded within the first portion of the temple. The second portion conceals the electrical connector from an exterior of the eyewear when the second portion connects to the first portion in a concealed state, and exposes the electrical connector from the exterior of the eyewear when disconnected from the first portion in an exposed state.
Methods and systems are disclosed for performing operations for providing a shared augmented reality experience in a video chat. A video chat can be established between a plurality of client devices. During the video chat, videos of users associated with the client devices can be displayed. During the video chat, a request from a first client device to activate a first AR experience can be received, and in response, and body parts of users depicted in the videos are modified to include one or more AR elements associated with the first AR experience.
H04L 12/18 - Arrangements for providing special services to substations for broadcast or conference
G06T 19/00 - Manipulating 3D models or images for computer graphics
H04L 51/52 - User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail for supporting social networking services
Examples in the present disclosure relate to the prediction of motion of a body part by an extended reality (XR) device. Tracking data is captured by one or more sensors associated with the XR device. The tracking data is processed to track the body part. Based on the tracking of the body part and a kinematic model of the body part, kinematic state tracking data is dynamically updated. The kinematic model and the kinematic state tracking data are used to generate a predicted future kinematic state of the body part. In some examples, operation of the XR device is controlled based on the predicted future kinematic state.
An eXtended Reality (XR) system provides grasp detection of a user grasping a virtual object. The grasp detection may be used as a user input into an XR application. The XR system provides a user interface of the XR application to a user of the XR system, the user interface including one or more virtual objects. The XR system captures video frame tracking data of a pose of a hand of a user while the user interacts with a virtual object of the one or more virtual objects and generates skeletal model data of the hand of the user based on the video frame tracking data. XR system generates grasp detection data based on the skeletal model data and virtual object data of the virtual object, and provides the grasp detection data to the XR application as user input into the XR application.
Augmented reality experiences with an eyewear device including a position detection system and a display system are provided. The eyewear device detects at least one of a hand gesture or movement of the user in the physical environment, and associates the hand gesture with a setting for a virtual object held within the memory of the eyewear device. The eyewear device may then change an attribute of the virtual object based on the detected hand gesture or movement. The eyewear device then provides an output corresponding to one or more attributes of the virtual object. The virtual object may be, for example, a music player or a virtual game piece.
G06F 3/04815 - Interaction with a metaphor-based environment or interaction object displayed as three-dimensional, e.g. changing the user viewpoint with respect to the environment or object
A method of correcting perspective distortion of a selfie image captured with a short camera-to-face distance by processing the selfie image and generating an undistorted selfie image appearing to be taken with a longer camera-to-face distance. A pre-trained 3D face GAN processes the selfie image, inverts the 3D face GAN to obtain improved face latent code and camera parameters, fine tunes a 3D face GAN generator, and manipulates camera parameters to render a photorealistic face selfie image. The processed selfie image has less distortion in the forehead, nose, cheek bones, jaw line, chin, lips, eyes, eyebrows, ears, hair, and neck of the face.
Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing at least one program, method, and user interface to facilitate a shared control of a virtual object by two or more users. A virtual object is displayed by a first device, for example, as part of an augmented reality experience where the display of the object is overlaid on a real-world environment. User input indicative of a modification to the virtual object is received. The virtual object is modified, and a modified virtual object is displayed by a second device.
The described system facilitates personalized avatar interactions by dynamically generating and displaying modified garments. The system determines the initiation of an interaction function by a first user with a second user within an interaction platform. It accesses avatar data for the first user, including visual attributes and a garment associated with the first avatar, as well as avatar data for the second user, including visual attributes and a garment associated with the second avatar. An image is generated featuring the first avatar wearing its garment alongside the second avatar wearing its garment. This image is applied to the first avatar's garment, creating a modified version that visually incorporates the second avatar. The system then displays the first avatar wearing the modified garment alongside the second avatar wearing its original garment, enabling enhanced user engagement through visually personalized and contextual representations of user interactions.
An artificial intelligence (AI) network or neural network is trained, using a relatively small number of reference images of a target garment, to enable virtual clothing try-ons of the target garment. Example methods include determining a pose for a person depicted in an input image, determining an area of the input image to replace with a target garment, changing values of pixels within the area, and inputting the pose, the area, and a text prompt describing the target garment, into a neural network, to generate an output image, wherein the neural network is trained to generate the target garment. Example methods include training the neural network with images of clothing in a same class or category as the target garment to teach the neural network to shape the target garment in accordance with a pose of the person and to preserve other clothing and the background.
A method for reducing motion-to-photon latency for hand tracking is described. In one aspect, a method includes accessing a first frame from a camera of an Augmented Reality (AR) device, tracking a first image of a hand in the first frame, rendering virtual content based on the tracking of the first image of the hand in the first frame, accessing a second frame from the camera before the rendering of the virtual content is completed, the second frame immediately following the first frame, tracking, using the computer vision engine of the AR device, a second image of the hand in the second frame, generating an annotation based on tracking the second image of the hand in the second frame, forming an annotated virtual content based on the annotation and the virtual content, and displaying the annotated virtual content in a display of the AR device.
Systems and methods are provided for presenting videos. The systems and methods access a video playback graphical user interface (GUI) that automatically plays back a plurality of videos in sequence. The systems and methods determine, by the one or more processors, a current mute state of the video playback GUI, a disabled mute state allowing output of audio associated with the playback of the plurality of videos, and an enabled mute state preventing the output of the audio associated with the playback of the plurality of videos. The systems and methods conditionally present an indicator that visually informs a user that audio is currently in the enabled mute state while an individual video of the plurality of videos is being played back based on the current mute state of the GUI.
H04N 21/439 - Processing of audio elementary streams
H04N 21/472 - End-user interface for requesting content, additional data or servicesEnd-user interface for interacting with content, e.g. for content reservation or setting reminders, for requesting event notification or for manipulating displayed content
35.
AUTOMATICALLY GENERATING DESCRIPTIONS OF AUGMENTED REALITY EFFECTS
A first image and a second image are accessed. The second image is generated by applying an augmented reality (AR) effect to the first image. The first image, the second image, and a prompt are provided to a visual-semantic machine learning model to obtain output describing at least one feature of the AR effect. A description of the AR effect is generated based on the output of the visual-semantic machine learning model. The description of the AR effect is stored in association with an identifier of the AR effect.
Methods and systems are disclosed for generating an extended reality (XR) try-on experience. The methods and systems store, in a multimodal memory, interaction data representing use of one or more interaction functions including data in different modalities. The methods and systems detect an object depicted in an image captured by an interaction client and generate, by a machine learning model, a prompt based on the object depicted in the image and the interaction data in the multimodal memory. The methods and systems generate an artificial texture based on the prompt and modify a texture of the object depicted in the image using the artificial texture that has been generated based on the prompt.
A system for providing context based creative tools and configured to perform operations that include: causing display of first media content at a client device, the first media content comprising a media attribute; detecting the media attribute of the first media content; presenting a graphical icon based on the media attribute of the media content at the client device; receiving an input that selects the graphical icon; and generating second media content at the client device based on the input that selects the graphical icon.
G06F 16/587 - Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using geographical or spatial information, e.g. location
38.
REDUCING POWER CONSUMPTION OF EXTENDED REALITY DEVICES
Examples describe a method performed by an extended reality (XR) device that implements a multi-camera object tracking system. The XR device accesses object tracking data associated with an object in a real-world environment. Based on the object tracking data, the XR device activates a low-power mode of the multi-camera object tracking system. In the low-power mode, a state of the object in the real-world environment is determined by using the multi-camera object tracking system.
Systems and methods in the present disclosure relate to the sharing of hand tracking data in the context of multi-user extended reality (XR) experiences. A first XR device of a first user and a second XR device of a second user participate in a shared XR experience. The first XR device receives, from the second XR device and via a communication link, hand tracking data for a second hand of the second user. The first XR device captures images that include a first hand of the first user and the second hand of the second user. While the shared XR experience is in progress, tracking operations performed by the first XR device are controlled based on the images captured by the first XR device and the hand tracking data it receives from the second XR device.
Described herein are systems and methods for optimizing neural network models for deployment on resource-constrained computing devices through layer-specific quantization. An original neural network model and deployment constraints are received as inputs. The optimization process alternates between a learning phase that updates model weights using task-specific loss functions and a compression phase that determines optimal bitwidth allocations for each layer through multiple-choice knapsack optimization. The compression phase computes quantization errors for different bitwidth options per layer and selects optimal bitwidth combinations while satisfying deployment constraints. The process iteratively updates a penalty parameter and continues until convergence, producing an optimized neural network model with quantized weights and layer-specific bitwidth allocations that maintains performance while meeting size, computational, and latency constraints for the target device.
Systems and methods in the present disclosure relate to the sharing of hand tracking data in the context of multi-user extended reality (XR) experiences. A first XR device of a first user and a second XR device of a second user participate in a shared XR experience. The first XR device receives, from the second XR device and via a communication link, hand tracking data for a second hand of the second user. The first XR device captures images that include a first hand of the first user and the second hand of the second user. While the shared XR experience is in progress, tracking operations performed by the first XR device are controlled based on the images captured by the first XR device and the hand tracking data it receives from the second XR device.
Systems and methods are provided for navigating messaging application interfaces. The systems and methods include operations for: displaying, by a messaging application of a user device, a menu comprising a first set of options relating to a first level in a hierarchy of levels; detecting, by a touch sensor, one finger touch of a first option of the first set of options; in response to detecting the one finger touch of the first option, displaying, by the messaging application, a second set of options related to the first option, the second set of options relating to a second level in the hierarchy of levels; detecting, by the touch sensor, two finger touch of a second option of the second set of options; and in response to detecting the two finger touch of the second option, re-displaying, by the messaging application, the first set of options.
G06F 3/0482 - Interaction with lists of selectable items, e.g. menus
G06F 3/04847 - Interaction techniques to control parameter settings, e.g. interaction with sliders or dials
G06F 3/0488 - Interaction techniques based on graphical user interfaces [GUI] using specific features provided by the input device, e.g. functions controlled by the rotation of a mouse with dual sensing arrangements, or of the nature of the input device, e.g. tap gestures based on pressure sensed by a digitiser using a touch-screen or digitiser, e.g. input of commands through traced gestures
43.
VIRTUAL EVALUATION TOOLS FOR AUGMENTED REALITY EXERCISE EXPERIENCES
Example systems, devices, media, and methods are described for evaluating movements and physical exercises in augmented reality using the display of an eyewear device. A motion evaluation application implements and controls the capturing of frames of motion data using an inertial measurement unit (IMU) on the eyewear device. The method includes presenting virtual targets on the display, localizing the current eyewear device location based on the captured motion data, and presenting virtual indicators on the display. The virtual targets represent goals or benchmarks for the user to achieve using body postures. The method includes detecting determining whether the eyewear device location represents an intersecting posture relative to the virtual targets, based on the IMU data. The virtual indicators display real-time feedback about user posture or performance relative to the virtual targets.
G06F 3/01 - Input arrangements or combined input and output arrangements for interaction between user and computer
G06F 3/0354 - Pointing devices displaced or positioned by the userAccessories therefor with detection of 2D relative movements between the device, or an operating part thereof, and a plane or surface, e.g. 2D mice, trackballs, pens or pucks
G06V 40/20 - Movements or behaviour, e.g. gesture recognition
Method of generating a real-time avatar animation starts with a processor receiving acoustic segments of a real-time acoustic signal. For each of the acoustic segments, processor generates using a music analyzer neural network a tempo value and a dance energy category and selects dance tracks based on the tempo value and the dance energy category. Processor generates using the dance tracks dance sequences for avatars, generates real-time animations for the avatars based on the dance sequences and avatar characteristics for the avatars, and causes to be displayed on a first client device the real-time animations of the avatars. Other embodiments are described herein.
Various embodiments provide for systems, methods, and computer-readable storage media for annotating a collection of media items, such as digital images. According to some embodiments, an annotation system automatically determines one or more annotations for a plurality of media content items, and generates a collection of media content items that associates the determined annotations with the plurality of media content items. Depending on the embodiment, annotations that may be determined for the plurality of media content (and associated with the collection for the media content items) can include, without limitation, a caption, a geographic location, a category, a novelty measurement, an event, and a highlight media content item representing the collection.
G06F 16/48 - Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
G06F 16/487 - Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually using geographical or spatial information, e.g. location
46.
INTRINSIC PARAMETERS ESTIMATION IN VISUAL TRACKING SYSTEMS
A method for adjusting camera intrinsic parameters of a multi-camera visual tracking device is described. In one aspect, a method for calibrating the multi-camera visual tracking system includes disabling a first camera of the multi-camera visual tracking system while a second camera of the multi-camera visual tracking system is enabled, detecting a first set of features in a first image generated by the first camera after detecting that the temperature of the first camera is within the threshold of the factory calibration temperature of the first camera, and accessing and correcting intrinsic parameters of the second camera based on the projection of the first set of features in the second image and a second set of features in the second image.
G06T 7/80 - Analysis of captured images to determine intrinsic or extrinsic camera parameters, i.e. camera calibration
G06V 10/44 - Local feature extraction by analysis of parts of the pattern, e.g. by detecting edges, contours, loops, corners, strokes or intersectionsConnectivity analysis, e.g. of connected components
H04N 17/00 - Diagnosis, testing or measuring for television systems or their details
H04N 23/61 - Control of cameras or camera modules based on recognised objects
H04N 23/65 - Control of camera operation in relation to power supply
H04N 23/68 - Control of cameras or camera modules for stable pick-up of the scene, e.g. compensating for camera body vibrations
47.
MODEL FINE-TUNING FOR AUTOMATED AUGMENTED REALITY DESCRIPTIONS
A second input image is generated by applying a target augmented reality (AR) effect to a first input image. The first input image and the second input image are provided to a first visual-semantic machine learning model to obtain output describing at least one feature of the target AR effect. The first visual-semantic machine learning model is fine-tuned from a second visual-semantic machine learning model by using training samples. Each training sample comprises a first training image, a second training image, and a training description of a given AR effect. The second training image is generated by applying the given AR effect to the first training image. A description of the target AR effect is selected based on the output of the visual-semantic machine learning model. The description of the target AR effect is stored in association with an identifier of the target AR effect.
Examples disclosed herein describe visual-inertial tracking techniques for extended reality (XR) devices. According to some example methods, an XR device is located in, and movable relative to, a vehicle. The XR device generates device tracking data and accesses vehicle tracking data. The vehicle tracking data is generated by an external sensor configured to measure motion of the vehicle. Consolidated tracking data is generated based on the device tracking data and the vehicle tracking data. In some examples, a pose of the XR device is determined by using the consolidated tracking data.
In various embodiments, a waveguide assembly includes first and second waveguide slabs. The first waveguide slab is to receive image-bearing light and enlarge a pupil size of the light parallel to a first axis. The first waveguide slab includes a first input-coupling device to couple image-bearing light into the first waveguide slab under total internal reflection (TIR), and a first out-coupling region to decouple image-bearing light out of the first waveguide slab by reflection. The second waveguide slab is to couple at least a portion of the out-coupled image-bearing light from the first waveguide slab into the second waveguide slab and enlarge the pupil size parallel to a second axis, which is substantially orthogonal to the first axis. The second waveguide slab includes a diffractive in-coupling region and a transmissive diffractive out-coupling region, through which a user can view real world imagery and the out-coupled image-bearing light simultaneously.
A waveguide for use in an augmented reality or virtual reality display comprises a plurality of optical structures in or on a photonic crystal. The plurality of optical structures are arranged in an array to provide two diffractive optical elements overlaid on one another in the waveguide. Each of the two diffractive optical elements is configured to receive light from an input direction and couple it towards the other diffractive optical element which can then act as an output diffractive optical element, providing outcoupled orders towards a viewer. The plurality of optical structures have different respective cross sectional shapes, for a cross section parallel to the plane of the waveguide, at different positions in the array in order to provide different diffraction efficiencies at different positions in the array.
H04M 1/72406 - User interfaces specially adapted for cordless or mobile telephones with means for local support of applications that increase the functionality by software upgrading or downloading
52.
AUGMENTED REALITY GAMING USING VIRTUAL EYEWEAR BEAMS
Interactive augmented reality experiences with an eyewear device including a virtual eyewear beam. The user can direct the virtual beam by orienting the eyewear device or the user’s eye gaze or both. The eyewear device may detect the direction of an opponent’s eyewear device or eye gaze of both. The eyewear device may calculate a score based on hits of the virtual beam of the user and the opponent on respective target areas such as the other player’s head or face.
A method to enhance virtual audio capture in Augmented Reality (AR) experience recordings starts with a processor receiving a video from a camera that includes images of a real-world scene and an AR content item. Processor receives acoustic signals from microphones generate acoustic signals using real-world audio and speaker output that including AR audio of the AR content item. Processor receives an audio file associated with the AR audio of the AR content item and generates an enhanced audio using the acoustic signals and the audio file. Processor generates an enhanced video using the video and the enhanced audio. Other examples are described herein.
A system is disclosed, including a processor and a memory. The memory stores instructions that, when executed by the processor, configure the system to perform operations. Raw region of interest (ROI) information is obtained, identifying a raw ROI having a raw ROI size, within a first video frame captured by a camera. Motion information representative of motion of the raw ROI relative to a field of view (FOV) of the camera is obtained. The raw ROI information and the motion information are processed to generate a dynamic ROI having a dynamic ROI size larger than the raw ROI size. A second video frame is captured by the camera. A portion of the second video frame defined by the dynamic ROI is processed to generate autoexposure (AE) settings for the camera.
An eXtended Reality (XR) system that determines a non-interaction intent of a user is provided. The XR system captures hand tracking data using tracking sensors that include cameras capable of capturing hand movements and gestures in real-time. The XR system also captures pose data of the head-wearable apparatus using pose sensors including an Inertial Measurement Unit (IMU) and cameras to determine Six Degrees of Freedom (6 DoF) data. The XR system detects non-interaction indicators by analyzing the hand tracking data to identify situations where the user likely does not intend to interact with virtual content. Based on these detected non-interaction indicators, the system modifies virtual interaction capabilities of an XR user interface such as by selectively enabling or disabling virtual cursor feedback while maintaining direct manipulation abilities.
The present disclosure provides a method and system for training a neural network to be executed by a processor system comprising a plurality of processor cores. During training, an aggregate loss is computed that is based on a task loss indicative of an accuracy with which the neural network performs a task and based on a computation performance loss which indicates an estimation of a computation performance parameter. Weights of the neural network are then updated by backpropagation of the aggregated loss.
Examples described herein relate to systems and methods for automatic evaluation of graphical element recommendations, such as sticker recommendations. According to some examples, a system accesses a set of text queries and provides each text query as input to a graphical element recommendation machine learning model. The graphical element recommendation machine learning model is trained to generate, based on a given text query, one or more graphical element recommendations for use in a message in a context of a messaging interface of an interaction application. The system obtains, from the graphical element recommendation machine learning model, at least one graphical element recommendation for each text query. The system generates a model quality score for the graphical element recommendation machine learning model by applying a model quality metric to the graphical element recommendations. Output indicative of the model quality score may be presented at a user device.
An extended Reality (XR) system that determines a non-interaction intent of a user is provided. The XR system captures hand tracking data using tracking sensors that include cameras capable of capturing hand movements and gestures in real-time. The XR system also captures pose data of the head-wearable apparatus using pose sensors including an Inertial Measurement Unit (IMU) and cameras to determine Six Degrees of Freedom (6D0F) data. The XR system detects non-interaction indicators by analyzing the hand tracking data to identify situations where the user likely does not intend to interact with virtual content. Based on these detected non-interaction indicators, the system modifies virtual interaction capabilities of an XR user interface such as by selectively enabling or disabling virtual cursor feedback while maintaining direct manipulation abilities.
A messaging system, which hosts a backend service for an associated messaging client, includes a voice chat system that provides voice chat functionality that enables users to dictate their messages, while delivering the resulting message to the intended recipient as both the associated audio and text content. When a user at a sender client device begins dictating a voice message, the voice chat system starts converting the received audio stream into text and, also, starts communicating the audio content together with the generated text to the recipient client device. The recipient user can listen to the voice message and read the text generated from the audio in real time. It is also possible for the recipient user to consume the voice message in a textual form only, if the sound at the client device is undesirable.
G10L 15/22 - Procedures used during a speech recognition process, e.g. man-machine dialog
G10L 15/30 - Distributed recognition, e.g. in client-server systems, for mobile phones or network applications
G10L 25/51 - Speech or voice analysis techniques not restricted to a single one of groups specially adapted for particular use for comparison or discrimination
G10L 25/90 - Pitch determination of speech signals
H04L 51/04 - Real-time or near real-time messaging, e.g. instant messaging [IM]
Systems and methods herein describe generating user-defined contextual spaces by receiving, by an image of a view of a physical space from a camera, receiving a virtual boundary of a first portion of the physical space, determining a virtual volume of the first portion of the physical space based on the first input, receiving data associated with to the first portion of the physical space, and storing the virtual volume in association with the received data.
A neural network processor is provided comprising a plurality of mutually succeeding neural network processor layers is provided. A neural network processor layer therein comprising a plurality of neural network processor elements (1) having a respective state register (2) for storing a state value (X) indicative for their state, as well as an additional state register (4) for storing a value (Q) of a state value change indicator that is indicative for a direction of a previous state change exceeding a threshold value. Neural network processor elements in a neural network processor layer are configured to selectively transmit differential event messages indicative for a change of their state, dependent both on the change of their state value and on the value of their state value change indicator.
G06N 3/04 - Architecture, e.g. interconnection topology
G06F 18/2113 - Selection of the most significant subset of features by ranking or filtering the set of features, e.g. using a measure of variance or of feature cross-correlation
At least one unit of a software application is identified. The at least one unit includes source code. The source code of the at least one unit is analyzed to determine a style of the source code. Metadata is extracted from the at least one unit based on the source code analysis. One or more features of the extracted metadata are classified. A template file is modified based on the extracted metadata and the classified features to create a modified template file.
A display driver device (210) receives a downloadable “sequence” for dynamically reconfiguring displayed image characteristics in an image system. The display driver device comprises one or more storage devices, for example, memory devices, for storing image data (218) and portions of drive sequences (219) that are downloadable and/or updated in real time depending on various inputs (214).
G09G 3/20 - Control arrangements or circuits, of interest only in connection with visual indicators other than cathode-ray tubes for presentation of an assembly of a number of characters, e.g. a page, by composing the assembly by combination of individual elements arranged in a matrix
09 - Scientific and electric apparatus and instruments
14 - Precious metals and their alloys; jewelry; time-keeping instruments
16 - Paper, cardboard and goods made from these materials
21 - HouseHold or kitchen utensils, containers and materials; glassware; porcelain; earthenware
25 - Clothing; footwear; headgear
28 - Games; toys; sports equipment
35 - Advertising and business services
41 - Education, entertainment, sporting and cultural services
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable computer software, namely, software for augmented reality (AR) and virtual reality (VR) for integrating electronic data with real-world environments and for creating, capturing, editing, and sharing multimedia content and data; downloadable software for creating, editing, storing, and transmitting digital content and communications via the internet, communication networks, and mobile devices; downloadable augmented reality software for virtual product try-on and visualization of consumer products; downloadable software for object and sound recognition and environmental scanning; downloadable artificial intelligence and machine learning software; computer hardware, computer peripherals, and wearable computer hardware and peripherals; computer hardware and peripherals for capturing, transmitting, and displaying images, video, audio, and data; downloadable software for setting up, configuring, and controlling wearable computer hardware and peripherals; downloadable multimedia files containing audio and video content in the fields of entertainment, photography, and social networking; computer software for accessing and transmitting data and content among electronic devices and displays Key chains Stationery; Writing pads; notebooks; stickers; printed greetings cards and printed postcards Containers; Mugs; drinking glasses; thermally insulated beverage containers Clothing; clothing, namely, shirts, sweatshirts, hats Toys and games; plush toys; beach balls; playing cards Retail store services featuring computer hardware, peripherals, cameras, video cameras, and digital media, namely, pre-recorded music, videos, photographs, images, and audiovisual content; facilitating the exchange and sale of goods and services of third parties via the internet and communication networks, namely, facilitating transactions between buyers and sellers through providing buyers with information about sellers, goods, and/or services via the internet and communication networks; advertising, marketing, and promotion services; dissemination of advertising for others via computer and other communication networks; consumer profiling for commercial or marketing purposes; providing commercial consumer information and advice for consumers in the selection of products to buy Entertainment services, namely, providing online non-downloadable multimedia content in the fields of entertainment, photography, and social networking in the nature of avatars, graphic icons, symbols, images representing individuals, fanciful designs, comics, comic series, phrases, and graphical depictions of people, places and things that end users can transmit and receive by means of the Internet or other computer or telecommunication networks, wireless communications networks, or by using computers, laptops, mobile equipment, and handheld digital electronic devices; providing online augmented reality experiences for entertainment purposes; organizing and hosting social entertainment events in the fields of augmented reality and creating, editing, sharing and sending digital photos, videos, images, emojis, avatars, emoticons, text, audio, and video games Providing online non-downloadable software for augmented reality (AR) and virtual reality (VR) for integrating electronic data with real-world environments and for creating, capturing, editing, and sharing multimedia content and data; providing online non-downloadable software for creating, editing, storing, and transmitting digital content and communications via the internet, communication networks, and mobile devices; providing online non-downloadable augmented reality software for virtual product try-on and visualization of consumer products; providing online non-downloadable software for object and sound recognition and environmental scanning; providing online non-downloadable artificial intelligence and machine learning software; providing online non-downloadable software for setting up, configuring, and controlling wearable computer hardware and peripherals; providing online non-downloadable software for accessing and transmitting data and content among electronic devices and displays; platform as a service (PaaS) featuring software for the foregoing purposes
09 - Scientific and electric apparatus and instruments
14 - Precious metals and their alloys; jewelry; time-keeping instruments
16 - Paper, cardboard and goods made from these materials
21 - HouseHold or kitchen utensils, containers and materials; glassware; porcelain; earthenware
25 - Clothing; footwear; headgear
28 - Games; toys; sports equipment
35 - Advertising and business services
41 - Education, entertainment, sporting and cultural services
42 - Scientific, technological and industrial services, research and design
Goods & Services
Downloadable computer software, namely, software for augmented reality (AR) and virtual reality (VR) for integrating electronic data with real-world environments and for creating, capturing, editing, and sharing multimedia content and data; downloadable software for creating, editing, storing, and transmitting digital content and communications via the internet, communication networks, and mobile devices; downloadable augmented reality software for virtual product try-on and visualization of consumer products; downloadable software for object and sound recognition and environmental scanning; downloadable artificial intelligence and machine learning software; computer hardware, computer peripherals, and wearable computer hardware and peripherals; computer hardware and peripherals for capturing, transmitting, and displaying images, video, audio, and data; downloadable software for setting up, configuring, and controlling wearable computer hardware and peripherals; downloadable multimedia files containing audio and video content in the fields of entertainment, photography, and social networking; computer software for accessing and transmitting data and content among electronic devices and displays Key chains Stationery; Writing pads; notebooks; stickers; printed greetings cards and printed postcards Containers; Mugs; drinking glasses; thermally insulated beverage containers Clothing; clothing, namely, shirts, sweatshirts, hats Toys and games; plush toys; beach balls; playing cards Retail store services featuring computer hardware, peripherals, cameras, video cameras, and digital media, namely, pre-recorded music, videos, photographs, images, and audiovisual content; facilitating the exchange and sale of goods and services of third parties via the internet and communication networks, namely, facilitating transactions between buyers and sellers through providing buyers with information about sellers, goods, and/or services via the internet and communication networks; advertising, marketing, and promotion services; dissemination of advertising for others via computer and other communication networks; consumer profiling for commercial or marketing purposes; providing commercial consumer information and advice for consumers in the selection of products to buy Entertainment services, namely, providing online non-downloadable multimedia content in the fields of entertainment, photography, and social networking in the nature of avatars, graphic icons, symbols, images representing individuals, fanciful designs, comics, comic series, phrases, and graphical depictions of people, places and things that end users can transmit and receive by means of the Internet or other computer or telecommunication networks, wireless communications networks, or by using computers, laptops, mobile equipment, and handheld digital electronic devices; providing online augmented reality experiences for entertainment purposes; organizing and hosting social entertainment events in the fields of augmented reality and creating, editing, sharing and sending digital photos, videos, images, emojis, avatars, emoticons, text, audio, and video games Providing online non-downloadable software for augmented reality (AR) and virtual reality (VR) for integrating electronic data with real-world environments and for creating, capturing, editing, and sharing multimedia content and data; providing online non-downloadable software for creating, editing, storing, and transmitting digital content and communications via the internet, communication networks, and mobile devices; providing online non-downloadable augmented reality software for virtual product try-on and visualization of consumer products; providing online non-downloadable software for object and sound recognition and environmental scanning; providing online non-downloadable artificial intelligence and machine learning software; providing online non-downloadable software for setting up, configuring, and controlling wearable computer hardware and peripherals; providing online non-downloadable software for accessing and transmitting data and content among electronic devices and displays; platform as a service (PaaS) featuring software for the foregoing purposes
66.
ENABLING THE VISUALLY IMPAIRED WITH AR USING FORCE FEEDBACK
A system and method provide feedback to a user, such as a visually impaired user, to guide the user to an object in the field of view of a camera mounted on a frame worn on the head of the user. A processor identifies at least one object and a body part of the user in the field of view of the camera and tracks relative positions of the body part relative to the identified object. The processor also generates and communicates at least one control signal for guiding the body part of the user to the identified object to a user feedback device worn on or adjacent the body part of the user. The feedback device receives the control signal(s) and converts the control signal(s) into at least one of sounds or haptic feedback that guides the body part to the identified object.
The technical problem of reducing the amount of processing involved when searching for customizable media content items that are suitable for incorporating input text is addressed by providing a hybrid search system. In some examples, the hybrid search system executes a rough search first, to determine whether a line of text can be incorporated into a media content item, based on character count conditions associated with the media content item. A more thorough evaluation of the input text with respect to the media content item is executed subsequent to the rough search if the rough search produces a result indicating uncertainty with respect to whether the combination of specific characters included in the input text can or cannot be incorporated into the media content item.
Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing a program and method for presenting participant reactions within a virtual working environment. The program and method provide a configuration interface for configuring a virtual working environment for plural participants, the configuration interface for specifying groups of participants, each group comprising respective participants selected from among the plural participants; receive first user input, provided via the configuration interface, specifying a first group of participants; provide, for each participant in the first group, display of a reactions interface with user-selectable buttons to indicate respective reactions for displaying to the first group; receive second user input, provided via the reactions interface, selecting one of the user-selectable buttons to indicate a reaction for displaying to the first group; and provide, for each participant in the first group, display of a reaction icon corresponding to the reaction.
G06F 3/04815 - Interaction with a metaphor-based environment or interaction object displayed as three-dimensional, e.g. changing the user viewpoint with respect to the environment or object
G06F 3/04817 - Interaction techniques based on graphical user interfaces [GUI] based on specific properties of the displayed interaction object or a metaphor-based environment, e.g. interaction with desktop elements like windows or icons, or assisted by a cursor's changing behaviour or appearance using icons
Aspects of the present disclosure involve a system for performing ray tracing between augmented reality (AR) and real-world objects. The system accesses, by the mobile device, a video depicting a first object. The system obtains, by the mobile device, a three-dimensional (3D) model of the first object. The system applies, by the mobile device, a ray tracing process to the 3D model of the first object to estimate an optical effect on a portion of the first object relative to a second object that is depicted in the video. The system modifies a visual property of the portion of the first object based on the optical effect relative to the second object.
Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing a program and method for providing make-up based augmented reality content. The program and method provide for receiving a request to present augmented reality content in association with a captured image depicting a face of the user; accessing an augmented reality content item associated with applying makeup to the face and configured to generate a mesh for tracking plural regions of the face; receiving user input selecting a region; determining at least one of a range of color values or a range of contrast values relating to available makeup products for the selected region; and presenting an interface element in association with the face, the interface element for user selection of at least one of a color value within the range of color values or a contrast value within the range of contrast values.
A45D 44/00 - Other cosmetic or toiletry articles, e.g. for hairdressers' rooms
G06F 3/04845 - Interaction techniques based on graphical user interfaces [GUI] for the control of specific functions or operations, e.g. selecting or manipulating an object, an image or a displayed text element, setting a parameter value or selecting a range for image manipulation, e.g. dragging, rotation, expansion or change of colour
G06F 3/04847 - Interaction techniques to control parameter settings, e.g. interaction with sliders or dials
G06F 3/04883 - Interaction techniques based on graphical user interfaces [GUI] using specific features provided by the input device, e.g. functions controlled by the rotation of a mouse with dual sensing arrangements, or of the nature of the input device, e.g. tap gestures based on pressure sensed by a digitiser using a touch-screen or digitiser, e.g. input of commands through traced gestures for inputting data by handwriting, e.g. gesture or text
A media player providing real time rewind playback of a played media file having segments of frames. A last segment N of the played media file is cached and rendered on a device, such as a mobile device, then a previous segment N-1 is cached and rendered, and the process continues until there are no more segments of the played media file to cache and render. Only a segment of the played media file is cached at a time, rather than the whole media file, such that the played media file can be replayed on the fly.
H04N 21/433 - Content storage operation, e.g. storage operation in response to a pause request or caching operations
G06F 3/04883 - Interaction techniques based on graphical user interfaces [GUI] using specific features provided by the input device, e.g. functions controlled by the rotation of a mouse with dual sensing arrangements, or of the nature of the input device, e.g. tap gestures based on pressure sensed by a digitiser using a touch-screen or digitiser, e.g. input of commands through traced gestures for inputting data by handwriting, e.g. gesture or text
H04N 19/177 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals using adaptive coding characterised by the coding unit, i.e. the structural portion or semantic portion of the video signal being the object or the subject of the adaptive coding the unit being a group of pictures [GOP]
H04N 19/426 - Methods or arrangements for coding, decoding, compressing or decompressing digital video signals characterised by implementation details or hardware specially adapted for video compression or decompression, e.g. dedicated software implementation characterised by memory arrangements using memory downsizing methods
72.
TECHNIQUES FOR GENERATING A STYLIZED MEDIA CONTENT ITEM WITH A GENERATIVE NEURAL NETWORK
A mobile application with an improved user interface facilitates generating stylized media content items including images and videos. An end-user selects a desired visual effect from a set of options. The mobile application captures or accesses an image. The image is processed on a server using a generative neural network pre-trained to apply stylizations based on the selected effect. The server sends back the stylized image to the mobile application for display. The end-user can then save the stylized image or generate a video (e.g., an animation) showing the original image transition to the stylized image. The user interface provides an efficient creative workflow to apply aesthetic enhancements in a visual style chosen by the end-user. Generative machine learning techniques automate stylization to enable accessible media customization and sharing.
G06T 11/60 - Editing figures and textCombining figures or text
G06F 3/0482 - Interaction with lists of selectable items, e.g. menus
G06F 3/0488 - Interaction techniques based on graphical user interfaces [GUI] using specific features provided by the input device, e.g. functions controlled by the rotation of a mouse with dual sensing arrangements, or of the nature of the input device, e.g. tap gestures based on pressure sensed by a digitiser using a touch-screen or digitiser, e.g. input of commands through traced gestures
A convolutional neural network processing system includes a data processor to process input feature map data and generate output feature map data in an output feature map storage space. Series of output feature data elements are stored at respective series of mutually successive locations in the output feature map storage space. The data processor identifies an input feature data element in the input feature map data, and accesses a series of mutually successive locations in the output feature map storage space. The data processor updates the output feature map data in an input-centric manner by using an update function to update a set of output feature data elements associated with the input feature data element by a convolution kernel of the update function. The set of output feature data elements is located in the accessed series of mutually successive locations.
A system and method are described for generating 3D garments from two-dimensional (2D) scribble images drawn by users. The system includes a conditional 2D generator, a conditional 3D generator, and two intermediate media including dimension-coupling color-density pairs and flat point clouds that bridge the gap between dimensions. Given a scribble image, the 2D generator synthesizes dimension-coupling color-density pairs including the RGB projection and density map from the front and rear views of the scribble image. A density-aware sampling algorithm converts the 2D dimension-coupling color-density pairs into a 3D flat point cloud representation, where the depth information is ignored. The 3D generator predicts the depth information from the flat point cloud. Dynamic variations per garment due to deformations resulting from a wearer's pose as well as irregular wrinkles and folds may be bypassed by taking advantage of 2D generative models to bridge the dimension gap in a non-parametric way.
Aspects of the present disclosure involve a system and a method for navigating images and AR experiences. The system and method present, by a messaging application, in a scrollable region on top of a live video feed being displayed in a GUI comprising a viewfinder, a first plurality of options associated with previously captured content items and a second plurality of options associated with AR experiences. The system and method scrolls the first plurality of options together with the second plurality of options of the scrollable region to bring one of the first plurality of options or one of the second plurality of options into focus. In response, the system and method: modify a configuration of the GUI in accordance with a first manner or modify the configuration of the GUI in accordance with a second manner.
G06F 3/04845 - Interaction techniques based on graphical user interfaces [GUI] for the control of specific functions or operations, e.g. selecting or manipulating an object, an image or a displayed text element, setting a parameter value or selecting a range for image manipulation, e.g. dragging, rotation, expansion or change of colour
G06F 3/04815 - Interaction with a metaphor-based environment or interaction object displayed as three-dimensional, e.g. changing the user viewpoint with respect to the environment or object
G06F 3/0482 - Interaction with lists of selectable items, e.g. menus
G06F 3/0488 - Interaction techniques based on graphical user interfaces [GUI] using specific features provided by the input device, e.g. functions controlled by the rotation of a mouse with dual sensing arrangements, or of the nature of the input device, e.g. tap gestures based on pressure sensed by a digitiser using a touch-screen or digitiser, e.g. input of commands through traced gestures
H04L 51/04 - Real-time or near real-time messaging, e.g. instant messaging [IM]
A system is disclosed, including a processor and a memory. The memory stores instructions that, when executed by the processor, configure the system to perform operations. Autoexposure (AE) primary camera information is obtained, identifying a first camera as an AE primary camera to be used for computing AE settings of the first camera and a second camera. First camera region of interest (ROI) information and second camera ROI information are obtained, representative of a first number of ROIs within a field of view (FOV) of the first camera and a second number of ROIs within a FOV of the second camera. In response to determining that the first number is zero and the second number is greater than zero, the AE primary camera information is updated to identify the second camera as the AE primary camera, and the second camera ROI information is processed to generate the AE settings.
The present disclosure seeks to address technical problems arising in the field of artificial intelligence (AI) by providing for training of a machine learning model to generate modified images based on an input image and an input instruction. For example, the machine learning model is trained to generate a modified portrait image based on an input portrait image and an input instruction. The machine learning model generates the modified portrait image to depict the input portrait image as modified according to the input instruction while maintaining the identity of a subject depicted in the input portrait image.
Eyewear providing an interactive augmented reality experience to allow a user of an eyewear device to display a 3D overlay image on a viewed person. The user can select the overlay image from a list of images, such as costumes, stored in memory or generated by the user. The images can be sorted in memory based on common attributes. Registration points of the person are continuously aligned with registration points of the overlay as the person moves such that the user appears to be wearing the 3D costume during movement. By aligning the registration points, the costume adapts to different body types and heights. The coloring of the costume can change based on the environment, such as the lighting, or to contrast with colors viewed in a viewfinder.
Described is a system for dynamically applying model adaptations customized for individual users by detecting an image of a first real-world object from a camera feed, detecting landmarks on the first real-world object, and processing the landmarks on the first real-world object using a generative machine learning model to generate a first custom image template for the first real-world object where portions of the first custom image template are populated with visual content placed based on the first custom image template. The system then applies a content augmentation based on the first custom image template to the camera feed.
The subject technology receives a set of inputs from multiple input sources. The subject technology determines a set of input features based on the set of inputs from the multiple input sources. The subject technology performs a time window-based aggregation on the set of input features to generate a set of aggregated features. The subject technology performs feature extraction, using a set of modular components of a modular classifier network, on the set of aggregated features to generate a set of extracted features. The subject technology generates, using a pinch detection head, a probability score indicating the likelihood of an occurrence of a pinch gesture based on the set of extracted features. The subject technology determines, using triggering logic, whether a pinch gesture has occurred based at least in part on the probability score. The subject technology provides a pinch detection output based at least in part on the determining.
G06F 3/01 - Input arrangements or combined input and output arrangements for interaction between user and computer
G06V 10/26 - Segmentation of patterns in the image fieldCutting or merging of image elements to establish the pattern region, e.g. clustering-based techniquesDetection of occlusion
G06V 10/40 - Extraction of image or video features
G06V 10/764 - Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
G06V 10/80 - Fusion, i.e. combining data from various sources at the sensor level, preprocessing level, feature extraction level or classification level
G06V 10/82 - Arrangements for image or video recognition or understanding using pattern recognition or machine learning using neural networks
G06V 40/20 - Movements or behaviour, e.g. gesture recognition
82.
AUGMENTED REALITY CONTENT GENERATORS FOR IDENTIFYING DESTINATION GEOLOCATIONS AND PLANNING TRAVEL
The subject technology causes display, at the client device, a set of augmented reality content items generated by the first augmented reality content generator. The subject technology receives, at the client device, a second selection of the particular augmented reality content item corresponding to the destination geolocation. The subject technology causes display, at the client device, a second set of augmented reality content items generated by the first augmented reality content generator. The subject technology receives, at the client device, a second selection of the second set of augmented reality content items. The subject technology causes display, at the client device, a third set of augmented reality content items generated by the first augmented reality content generator, the third set of augmented reality content items comprising at least one activity or location associated with the destination geolocation and a selected period of time.
Provided are systems and methods for providing personalized videos featuring multiple persons. An example method includes receiving a video including a plurality of frames including at least one target face, extracting target face parameters associated with the at least one target face, where the target face parameters include facial identity parameters and facial expression parameters, storing the target face parameters as metadata associated with at least one frame of the plurality of frames, receiving an image of a source face, generating source face parameters based on the image of the source face, generating an output face by combining the source face parameters with the facial expression parameters obtained from the metadata associated with the at least one frame, and generating a personalized video by replacing the at least one target face with the output face at least in the at least one frame.
Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing a program and method for configuring a three-dimensional (3D) model within a virtual conferencing system. The program and method provide, in association with designing a room for virtual conferencing, an interface for configuring a 3D model; receiving, via the interface, an indication of user input for setting properties for the 3D model, the properties specifying image data for projecting onto the 3D model; and in association with virtual conferencing, providing display of the room based on the properties for the 3D model, and causing the image data to be projected onto the 3D model within the room.
G06F 3/04815 - Interaction with a metaphor-based environment or interaction object displayed as three-dimensional, e.g. changing the user viewpoint with respect to the environment or object
G06F 3/04845 - Interaction techniques based on graphical user interfaces [GUI] for the control of specific functions or operations, e.g. selecting or manipulating an object, an image or a displayed text element, setting a parameter value or selecting a range for image manipulation, e.g. dragging, rotation, expansion or change of colour
G06F 3/04847 - Interaction techniques to control parameter settings, e.g. interaction with sliders or dials
G06T 19/00 - Manipulating 3D models or images for computer graphics
G06T 19/20 - Editing of 3D images, e.g. changing shapes or colours, aligning objects or positioning parts
Disclosed is a method of providing a music creation interface using a head-mounted device, including displaying first and second geometric loops fixed relative to a location in the real world, the first and second geometric loops each including a plurality of beat indicators. The second geometric loop is spaced apart from the first geometric loop. An interface comprising a plurality of sound or note icons is displayed, and in response to receiving user selection to move a selected sound or note icon to a particular beat indicator on one of the geometric loops, the selected sound or note icon is displayed at the particular beat indicator. In use, the geometric loops are rotated relative to at least one play indicator, and the selected sound or note icon is rendered when it reaches the at least one play indicator.
The subject technology requests a group identifier (ID) based on an item identification indicator to an extension application programming interface (API). The subject technology requests, using the extension API, a set of augmented reality (AR) content generator IDs, based on the item identification indicator, to a shopping AR content generator service. The subject technology receives the set of AR content generator IDs based on a mapping of the item identification indicator to a collection of AR content generator IDs. The subject technology requests, using the extension API, first metadata associated with the set of AR content generators and the item identification indicator to the shopping AR content generator service. The subject technology requests, using the extension API, a set of products based on the first metadata. The subject technology generates, using the extension API, a group ID based on the set of products. The subject technology receives the group ID.
Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing a program and a method for performing operations comprising: receiving, from a client device of a first user, a request from the first user to engage in an AR shopping experience curated by a store; identifying a first real-world product available for purchase from the store; receiving an image of a real-world environment of the first user; generating a first AR item that represents the first real-world product; comparing visual attributes of the first AR item to physical layouts of a plurality of real-world objects depicted in the image of the real-world environment; and overlaying the first AR item on a first real-world object of the plurality of real-world objects in the image responsive to comparing the visual attributes of the first AR item to the physical layouts of the plurality of real-world objects.
G06K 7/14 - Methods or arrangements for sensing record carriers by electromagnetic radiation, e.g. optical sensingMethods or arrangements for sensing record carriers by corpuscular radiation using light without selection of wavelength, e.g. sensing reflected white light
G06Q 10/087 - Inventory or stock management, e.g. order filling, procurement or balancing against orders
A system and method for generating augmented reality (AR) experiences are disclosed. The system generates source and target indications associated with an image transformation, and generates a first set of source images and first set of target images using a first trained machine learning (ML) model, the source indications, and the target indications. The system trains a second ML model to generate a target image corresponding to a source image based on the first set of source images and the first set of target images, and generates a second set of target images using the second trained ML model and a second set of source images. The system trains a third ML model to generate an additional target image corresponding to an additional source image based on the second set of source images and second set of target images, and generates an AR experience comprising the third trained ML model.
Systems and methods are presented for capturing a video in real-time by an image capture device using a skeletal pose system. The skeletal pose system identifies first pose information in the video, applies a first virtual effect to the video in response to identifying the first pose information, identifies second pose information in the video, and applies a second virtual effect to the video in response to identifying the first pose information.
Apparatuses, systems for electronic wearable devices such as smart glasses are described. The wearable device can comprise a frame, an elongate temple and an articulated joint. The frame can define one or more optical element holders configured to hold respective optical elements for viewing by a user in a viewing direction. The temple can be moveably connected to the frame for holding the frame in position when the device is worn by the user. The articulated joint can connect the temple and the frame to permit movement of the temple relative to the frame between a wearable position in which the temple is generally aligned with the viewing direction, and a collapsed position in which the temple extends generally transversely to the viewing direction. The articulated joint can include a base foot fixed to the frame and oriented transversely to the viewing direction.
A method for managing power resource in an augmented reality (AR) device is described. In one aspect, the method includes configuring a low-power mode to run on a low-power processor of the AR device using a first set of sensor data, and a high-power mode to run on a high-power processor of the AR device using a second set of sensor data, operating, using the low-power processor, a low-power application in the low-power mode based on the first set of sensor data, detecting a request to operate a high-power application at the AR device, in response to detecting the request, activating the second set of sensors of the AR device corresponding to the high-power mode, and operating, using the high-power processor, a high-power application in the high-power mode based on the second set of sensors.
Examples include a wearable device such as smart glasses having a frame, a temple and onboard electronics components. The frame can define one or more optical element holders for holding respective optical elements within view of a user when the eyewear body is worn. The pair of temples can be connected to the eyewear frame for supporting the eyewear frame in position within view of the user when the eyewear body is worn. The antenna can be incorporated in at least a first of the pair of temples. The antenna can include an exciter circuit, at least a portion of a battery flex of the first of the pair of temples configured as an active element of the antenna and a metal component of the first of the pair of temples configured as a ground of the antenna.
Methods, systems, mobile devices, and non-transitory computer-readable mediums for easily aesthetically enhancing images such as selfies. An example algorithm's input has three parts: image, manipulation magnitude, and text guidance. The algorithm includes two parts: (1) guidance generation based on public and personal aesthetic preferences, and (2) selfie generation. The first part outputs an image to maximize an aesthetic enhancement score (e.g., a beauty score) while following the manipulation input where the output image contains a manipulation direction. The second part is a conditional diffusion model that accepts the rendered output image from the first part and is conditioned on the input image and outputs the final image. The second part is personalized by the user's images.
Hierarchical patch-wise diffusion models (HPDMs) use a diffusion paradigm that learns a hierarchical distribution of patches instead of whole videos for efficient patch-wise training of diffusion models. To enforce consistency between the patches, deep context fusion may be used to propagate the context information from low-scale to high-scale patches in a hierarchical manner. To accelerate patch-wise training and inference, adaptive computation also may be used to allocate more computational resources and network capacity towards coarse image details and to cheapen synthesis of high-frequency texture details. All the processing stages are jointly trained to provide spatially aligned global context to the higher levels of the cascade. As a result, the model does not operate on the full-resolution inputs, which allows the model to be trained on high-resolution video datasets in an end-to-end fashion.
A messaging system performs image processing to relight objects with neural networks for images provided by users of the messaging system. A method of relighting objects with neural networks includes receiving an input image with first lighting properties comprising an object with second lighting properties and processing the input image using a convolutional neural network to generate an output image with the first lighting properties and comprising the object with third lighting properties, where the convolutional neural network is trained to modify the second lighting properties to be consistent with lighting conditions indicated by the first lighting properties to generate the third lighting properties. The method further includes modifying the second lighting properties of the object to generate the object with modified second lighting properties and blending the third lighting properties with the modified second lighting properties to generate a modified output image comprising the object with fourth lighting properties.
Described is a system for emphasizing XR content based on user intent by gathering interaction data from use of one or more interaction functions by a user, accessing a camera feed of a camera system from the XR device, analyzing a combination of data corresponding to the interaction data and the camera feed using a first machine learning model to identify a priority for individual media content items, and determining that a first subset of media content items are of a higher priority than a second subset of media content items. Then the system displays the media content items on the XR device of the user, the first subset of the media content items displayed differently than the second subset of the media content items based on the identified priority.
Described herein are techniques for facilitating the communication of text-based messages between end-users who are using messaging applications executing on client-based computing devices with different capabilities. Specifically, the messaging system described herein enables a first end-user to add a message element to a text-based message, which, when received by a message recipient using an augmented reality messaging application, will cause a 3-D avatar representing the message sender, to animate in accordance with a specific avatar animation associated with the message element. The message element may be an emoji, or a special sequence of characters, and may be a visible or invisible (e.g., meta-data) element of the text-based message.
G06T 19/00 - Manipulating 3D models or images for computer graphics
G06T 13/40 - 3D [Three Dimensional] animation of characters, e.g. humans, animals or virtual beings
G10L 13/08 - Text analysis or generation of parameters for speech synthesis out of text, e.g. grapheme to phoneme translation, prosody generation or stress or intonation determination
Systems, methods, and computer readable media that schedules requests for location data of a mobile device, where the methods include selecting a first positioning system based on a power requirement, a latency requirement, and an accuracy requirement, and determining whether a first condition is satisfied for querying the first positioning system. The method further comprises in response to a determination that the first condition is satisfied, querying the first positioning system for first position data. The method further comprises in response to a determination that the first condition is not satisfied, selecting a second positioning system based on the power requirement, the latency requirement, and the accuracy requirement, determining whether a second condition is satisfied for querying the second positioning system, and in response to a determination that the second condition is satisfied, querying the second positioning system for second position data.
G01S 5/00 - Position-fixing by co-ordinating two or more direction or position-line determinationsPosition-fixing by co-ordinating two or more distance determinations
G01S 5/02 - Position-fixing by co-ordinating two or more direction or position-line determinationsPosition-fixing by co-ordinating two or more distance determinations using radio waves
G01S 19/48 - Determining position by combining or switching between position solutions derived from the satellite radio beacon positioning system and position solutions derived from a further system
An optical waveguide device for use in a head up display. The waveguide device provides pupil expansion in two dimensions. The waveguide device comprise a primary waveguide and a secondary waveguide, the secondary waveguide being positioned on a face of the primary waveguide. The secondary waveguide has a diffraction grating on a face opposite to the face which contacts the primary waveguide. The diffraction grating diffracts light into more than diffraction order. Rays diffracted into a non-zero order are trapped in the secondary waveguide by total internal reflection.