AlexIndustrial/mesa

Author	SHA1	Message	Date
Ilia Mirkin	f96f210239	a2xx: only update rasterizer settings when they're there The rasterizer being empty can happen e.g. during clears Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2017-08-15 22:54:40 -04:00
Ilia Mirkin	08f72a8944	a2xx: add logicop support This passes both gl-1.0-logicop and gl-1.1-xor piglits. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2017-08-15 22:54:40 -04:00
Ilia Mirkin	978c4c597a	glsl/ast: update rhs in addition to the var's constant_value We continue in the code to do some more things with the rhs, including setting a constant initializer. If the type is wrong, this causes some confusion down the line, leading to assertions. This makes sure that the rhs processing continues to flow as-if the type was correct to start with (even though the state has been marked as an error state). Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=101766 Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org> Cc: mesa-stable@lists.freedesktop.org	2017-08-15 22:14:05 -04:00
Jason Ekstrand	98983503cb	anv: Advertise VK_KHR_external_semaphore Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-08-15 19:08:26 -07:00
Jason Ekstrand	55bce22d8d	anv: Use DRM sync objects for external semaphores when available Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-08-15 19:08:26 -07:00
Jason Ekstrand	f41a0e4b0d	anv/gem: Add a drm syncobj support Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-08-15 19:08:26 -07:00
Jason Ekstrand	5c4e4932e0	anv: Implement support for exporting semaphores as FENCE_FD Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-08-15 19:08:26 -07:00
Jason Ekstrand	e4054ab77b	anv/gem: Use EXECBUFFER2_WR when the FENCE_OUT flag is set Reviewed-by: Chad Versace <chadversary@chromium.org> Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-08-15 19:08:26 -07:00
Jason Ekstrand	017cdb10cf	anv: Submit a dummy batch when only semaphores are provided. Vulkan allows you to do a submit whose only job is to wait on and trigger semaphores. The easiest way for us to support that right now is to insert a dummy execbuf. Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-08-15 19:08:26 -07:00
Jason Ekstrand	031f57eba3	anv: Add a basic implementation of VK_KHX_external_semaphore This patch adds an implementation based on DRM BOs. We don't actually advertise the extension yet because we want to add a couple more paths first. Reviewed-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-08-15 19:08:26 -07:00
Aaron Watry	a8296dbd5a	clover/event: Include additional event statuses for clSetEventCallback From CL 2.0 Section 5.11 (Event Objects): clSetEventCallback returns CL_SUCCESS if the function is executed successfully. Otherwise, it returns one of the following errors: ... CL_INVALID_VALUE if pfn_event_notify is NULL or if command_exec_callback_type is not CL_SUBMITTED , CL_RUNNING or CL_COMPLETE . Fixes: OpenCL CTS test_conformance/events/test_events callbacks Signed-off-by: Aaron Watry <awatry@gmail.com> Reviewed-by: Francisco Jerez <currojerez@riseup.net>	2017-08-15 19:55:15 -05:00
Jonas Pfeil	494f86bbe5	broadcom/vc4: Port NEON-code to ARM64 Changed all register and instruction names, works the same. v2: Rebase on build system changes (by anholt) v3: Fix build on clang (by anholt, reported by Rob) Signed-off-by: Jonas Pfeil <pfeiljonas@gmx.de> Tested-by: Rob Herring <robh@kernel.org>	2017-08-15 13:23:54 -07:00
Eric Anholt	bd5efbd70b	broadcom/vc4: Build the vc4_tiling_lt_neon.c with -mfpu=neon on ARM. If you don't pass this, the compiler refuses to compile the assembly for pre-v7 CPUs. This also keeps us from building identical, non-NEON code on aarch64 and x86. Fixes: `a373f77662` ("vc4: Use a wrapper file to set VC4_BUILD_NEON instead of CFLAGS.") v2: Fix Android build by just appending NEON_C_SOURCES when ARCH_ARM_HAVE_NEON. Tested-by: Rob Herring <robh@kernel.org>	2017-08-15 13:23:54 -07:00
Eric Anholt	b94ddc181b	util: Fix build on old glibc. We need to link librt for u_thread.h's clock_gettime() call. Fixes: `b822d9dd67` ("gallium/util: move u_queue.{c,h} to src/util") Reviewed-by: Matt Turner <mattst88@gmail.com>	2017-08-15 13:23:54 -07:00
Eric Anholt	f785db3d31	broadcom: Add v3d_xml.h to gitignore.	2017-08-15 13:23:54 -07:00
Eric Anholt	463de32b95	broadcom: Add missing libexpat cflags for the decoder. The Raspbian ARMv6 cross compiler wasn't picking up my (amd64) system copy of the header the way that the system gcc and armhf cross-compile did.	2017-08-15 13:23:54 -07:00
Dave Airlie	694d59fbaf	radv/gfx9: for fast clear use is_linear flag. The legacy test won't work on gfx9. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Cc: "17.2" <mesa-stable@lists.freedesktop.org> Signed-off-by: Dave Airlie <airlied@redhat.com>	2017-08-16 06:27:30 +10:00
David Airlie	31bb8517a1	radv/gfx9: fix tile swizzle handling for gfx9 This sets the tile swizzle up properly for gfx9. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Cc: "17.2" <mesa-stable@lists.freedesktop.org> Signed-off-by: Dave Airlie <airlied@redhat.com>	2017-08-16 05:54:19 +10:00
David Airlie	e43cc3e3af	radv/gfx9: handle GFX9 opaque metadata port the opaque metadata changes from radeonsi for gfx9. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Cc: "17.2" <mesa-stable@lists.freedesktop.org> Signed-off-by: Dave Airlie <airlied@redhat.com>	2017-08-16 05:54:15 +10:00
David Airlie	674ecbfef2	radv: emit db_htile_surface reg on gfx9 as well This is also a GFX9 register. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Cc: "17.2" <mesa-stable@lists.freedesktop.org> Signed-off-by: Dave Airlie <airlied@redhat.com>	2017-08-16 05:54:09 +10:00
Dave Airlie	fc600eb98d	radv/gfx9: remove some leftover gfx6 descriptor setup. We set this later in the non-gfx9 path, just remove these bits from here. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Cc: "17.2" <mesa-stable@lists.freedesktop.org> Signed-off-by: Dave Airlie <airlied@redhat.com>	2017-08-16 05:54:03 +10:00
Dave Airlie	5247b311e9	radv/gfx9: fix set predication packet. The predication packet changed format on GFX9, update the driver. Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Cc: "17.2" <mesa-stable@lists.freedesktop.org> Signed-off-by: Dave Airlie <airlied@redhat.com>	2017-08-16 05:52:50 +10:00
Scott D Phillips	d6539608a4	intel/genxml: Fix gen10 BLEND_STATE variable length packing BLEND_STATE packing was modified to be variable-length in: `9670124e31` genxml: Make BLEND_STATE command support variable length array. The initial gen10.xml still had the old, fixed-length style definition for BLEND_STATE. So gen10_upload_blend_state would overwrite the packed BLEND_STATE_ENTRYs with its own fixed array of all-zero entries when packing BLEND_STATE. This caused BLEND_STATE upload to not work at all. Fixes: `aa416f515a` ("i965/genxml: Add gen10.xml") Reviewed-by: Rafael Antognolli <rafael.antognolli@intel.com> Reviewed-by: Anuj Phogat <anuj.phogat@gmail.com>	2017-08-15 09:06:29 -07:00
Timothy Arceri	fe74c8ffbf	mesa: count uniform against storage when its bindless Gallium drivers use this code path so we need to account for bindless after all. Fixes: `365d34540f` ("mesa: correctly calculate the storage offset for i915") Reviewed-by: Marek Olšák <marek.olsak@amd.com>	2017-08-15 23:51:35 +10:00
Marek Olšák	1ab7fed707	radeonsi: disable CE by default It makes performance worse by a very small (hard to measure) amount. We've done extensive profiling of this feature internally. Cc: 17.1 17.2 <mesa-stable@lists.freedesktop.org> Reviewed-by: Bas Nieuwenhuizen <bas@basnieuwenhuizen.nl> Acked-by: Christian König <christian.koenig@amd.com>	2017-08-15 15:03:43 +02:00
Dave Airlie	e0edfadec8	radeonsi: initialise imported surface to 0. For memobj imports we weren't setting the surface to 0, which meant sometimes we'd end up with tile_swizzle garbage, which would corrupt rendering. This seems to fix the image corruption on the imported memory objects in vrdashboard for me. Reviewed-by: Marek Olšák <marek.olsak@amd.com> Signed-off-by: Dave Airlie <airlied@redhat.com>	2017-08-15 01:35:58 +01:00
Timothy Arceri	de0e62e106	st/mesa: correctly calculate the storage offset When generating the storage offset for struct members we need to skip opaque types as they no longer have backing storage. Fixes: `fcbb93e860` ("mesa: stop assigning unused storage for non-bindless opaque types") Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=101983 Reviewed-by: Dave Airlie <airlied@redhat.com>	2017-08-15 08:20:57 +10:00
Timothy Arceri	365d34540f	mesa: correctly calculate the storage offset for i915 When generating the storage offset for struct members we need to skip opaque types as they no longer have backing storage. Fixes: `fcbb93e860` ("mesa: stop assigning unused storage for non-bindless opaque types") V2: simplify since bindless will never be supported in this code Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=101983 Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2017-08-15 08:20:57 +10:00
Ben Widawsky	1efd73df39	i965: Advertise the CCS modifier v2: Rename modifier to be more smart (Jason) FINISHME: Use the kernel's final choice for the fb modifier bwidawsk@norris2:~/intel-gfx/kmscube (modifiers $) ~/scripts/measure_bandwidth.sh ./kmscube none Read bandwidth: 603.91 MiB/s Write bandwidth: 615.28 MiB/s bwidawsk@norris2:~/intel-gfx/kmscube (modifiers $) ~/scripts/measure_bandwidth.sh ./kmscube ytile Read bandwidth: 571.13 MiB/s Write bandwidth: 555.51 MiB/s bwidawsk@norris2:~/intel-gfx/kmscube (modifiers $) ~/scripts/measure_bandwidth.sh ./kmscube ccs Read bandwidth: 259.34 MiB/s Write bandwidth: 337.83 MiB/s v2: Move all references to the new fourcc code(s) to this patch. v3: Rebase, remove Yf_CCS (Daniel) Signed-off-by: Ben Widawsky <ben@bwidawsk.net> Signed-off-by: Jason Ekstrand <jason@jlekstrand.net> Acked-by: Daniel Stone <daniels@collabora.com> Reviewed-by: Chad Versace <chadversary@chromium.org>	2017-08-14 10:43:30 -07:00
Jason Ekstrand	51600b8489	i965/miptree: More conservatively resolve external images Instead of always doing a full resolve, only resolve the bits that are needed. This means that we only do a partial resolve when the miptree modifier is I915_FORMAT_MOD_Y_TILED_CCS. Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com> Reviewed-by: Chad Versace <chadversary@chromium.org>	2017-08-14 10:43:30 -07:00
Ben Widawsky	8f6e54c929	i965: Pretend that CCS modified images are two planes v2: move is_aux into if block. (Jason) Use else block instead of goto (Jason) v3: Fix up logic for is_aux (Ben) Fix up size calculations and add FIXME (Ben) v4 (Jason Ekstrand): Use the aux_pitch in the image instead of calculating it Signed-off-by: Ben Widawsky <ben@bwidawsk.net> Acked-by: Daniel Stone <daniels@collabora.com> Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com> Reviewed-by: Chad Versace <chadversary@chromium.org>	2017-08-14 10:43:30 -07:00
Jason Ekstrand	a1e5db9888	i965/screen: Support import and export of surfaces with CCS Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com> Reviewed-by: Chad Versace <chadversary@chromium.org>	2017-08-14 10:43:30 -07:00
Ben Widawsky	a068fdc861	i965/miptree: Allocate mcs_buf for an image's CCS This code will disable actually creating these buffers for the scanout, but it puts the allocation in place. Primarily this patch is split out for review, it can be squashed in later if preferred. v2: assert(mt->offset == 0) in ccs creation (as requested by Topi) Remove bogus is_scanout check in miptree_release v3: Remove is_scanout assert in intel_miptree_create. It doesn't work with latest codebase - not sure it ever should have worked. v4: assert(mt->last_level == 0) and assert(mt->first_level == 0) in ccs setup (Topi) v5 (Jason Ekstrand): - Base the decision to allocate a CCS on the image modifier Signed-off-by: Ben Widawsky <ben@bwidawsk.net> Acked-by: Daniel Stone <daniels@collabora.com> Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com> Reviewed-by: Chad Versace <chadversary@chromium.org>	2017-08-14 10:43:30 -07:00
Ben Widawsky	f6fbeaf1c4	i965: Support images with aux buffers Previously images did not support any auxiliary compression surfaces (CCS, MCS, or HiZ). That's about to change. This patch just adds the fields to __DRIimageRec to make auxiliary surfaces possible. v2 (Jason Ekstrand): - Add an aux_pitch parameter as well as aux_offset Signed-off-by: Ben Widawsky <ben@bwidawsk.net> Acked-by: Daniel Stone <daniels@collabora.com> Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com> Reviewed-by: Chad Versace <chadversary@chromium.org>	2017-08-14 10:43:30 -07:00
Jason Ekstrand	cf2e92262b	intel/isl: Add support for I915_FORMAT_MOD_Y_TILED_CCS Reviewed-by: Topi Pohjolainen <topi.pohjolainen@intel.com> Reviewed-by: Chad Versace <chadversary@chromium.org>	2017-08-14 10:43:30 -07:00
Jason Ekstrand	51eb40d414	i965/screen: Stop redefining DRM_FORMAT_MOD_(INVALID\|LINEAR) Reviewed-by: Ben Widawsky <ben@bwidawsk.net>	2017-08-14 10:43:30 -07:00
Scott D Phillips	f7dfc44c61	i965/blorp: Correct type of src_format in call to intel_miptree_texture_aux_usage intel_miptree_texture_aux_usage() takes an isl_format, but we are passing a mesa_format. clang warns: brw_blorp.c:305:52: warning: implicit conversion from enumeration type 'mesa_format' to different enumeration type 'enum isl_format' [-Wenum-conversion] intel_miptree_texture_aux_usage(brw, src_mt, src_format); ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ^~~~~~~~~~ Fixes: `fc1639e46d` ("i965/blorp: Use texture/render_aux_usage for blits") Cc: "17.2" <mesa-stable@lists.freedesktop.org> Reviewed-by: Jordan Justen <jordan.l.justen@intel.com> Reviewed-by: Jason Ekstrand <jason@jlekstrand.net>	2017-08-14 10:41:54 -07:00
Julien Isorce	91d93aa621	st/va: change frame_idx from array to hash table The picture_id was assumed to be a frame number so in 0-31. But the vaapi client gstreamer-vaapi uses the surfaces handles as identifier which are unsigned int. This bug can happen when using a lot of vaapi surfaces within the same process. Indeed Mesa/st/va increments a counter for the surface ID: mesa/util/u_handle_table.c::handle_table_add which starts from 0 and incremented by 1 at each call. So creating more than 32 surfaces was a problem. The following bug contains a test that reproduces the problem by running a couple of vaapih264enc in the same process. The above also explains why there was no pb when running them in separated processes. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=102006 Signed-off-by: Julien Isorce <jisorce@oblong.com> Tested-by: Tomas Rataj <rataj28@gmail.com> Acked-by: Christian König <christian.koenig@amd.com> Reviewed-and-tested-by: Boyuan Zhang <Boyuan.Zhang@amd.com>	2017-08-14 13:40:19 +01:00
Ilia Mirkin	165e18dd21	nv50/ir: clean up saturated values immediately Since we don't iterate to a fixed point, we can end up in situations where we have a SAT instruction + a long immediate. This is not legal. However since it's immediately computable, just run unary straight away to handle the situation. Fixes: `24a799ad35` ("nv50/ir: fix ConstantFolding with saturation") Reported-by: Tobias Klausmann <tobias.johannes.klausmann@mni.thm.de> Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu> Cc: mesa-stable@lists.freedesktop.org	2017-08-12 14:49:08 -04:00
Ilia Mirkin	ea22ac23e0	nvc0/ir: unlink values pre- and post-call to division function While technically correct, this can lead to e.g. getImmediate assuming that it can walk up the value chain. It could be fixed to not do this, but it seems easier and less error-prone to just not link the two values to save on one LValue object. Signed-off-by: Ilia Mirkin <imirkin@alum.mit.edu>	2017-08-12 14:49:08 -04:00
Kenneth Graunke	22e1d8832c	i965: Guard GetBufferSubData's streaming memcpy load with USE_SSE41 This should hopefully fix build issues on 32-bit Android-x86. v2: s/USE_SSE4_1/USE_SS41/, caught by Gražvydas Ignotas. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=102050 Reviewed-by: Tapani Pälli <tapani.palli@intel.com> Reviewed-by: Emil Velikov <emil.velikov@collabora.com>	2017-08-12 01:42:32 -07:00
Kenneth Graunke	da0840246f	i965: Clean up intel_batchbuffer_init(). Passing screen lets us get the kernel features, devinfo, and bufmgr, without needing container_of. This use of container_of could cause crashes due to issues with the "sample" macro parameter. Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=102062 Reviewed-by: Tapani Pälli <tapani.palli@intel.com> Reviewed-by: Emil Velikov <emil.velikov@collabora.com> Reviewed-by: Iago Toral Quiroga <itoral@igalia.com>	2017-08-12 01:41:24 -07:00
Marek Olšák	b420680ede	gallium/radeon: only pass shader-specific debug flags to the disk shader cache Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com> Reviewed-by: Timothy Arceri <tarceri@itsqueeze.com>	2017-08-11 20:38:29 +02:00
Marek Olšák	d1285a7103	radeonsi/gfx9: fix the scissor bug workaround otherwise there is corruption in most apps. Fixes: `0fe0320` radeonsi: use optimal packet order when doing a pipeline sync Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2017-08-11 20:38:29 +02:00
Marek Olšák	27fef5d52d	radeonsi/gfx9: use the VI codepath for clamping Z This fixes corrupted shadows in Unigine Valley. The corruption disappeared when I stopped setting IMG_DATA_FORMAT_24_8 for depth. Cc: 17.2 <mesa-stable@lists.freedesktop.org> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2017-08-11 20:38:29 +02:00
Daniel Stone	2eee03b7a1	egl: Update headers from Khronos Taken from egl-registry 7d68647c4dab. Signed-off-by: Daniel Stone <daniels@collabora.com>	2017-08-11 11:16:00 +01:00
Daniel Stone	7d26a52a7a	egl/dri2: Allow modifiers to add FDs to imports When using dmabuf import, make sure that the modifier is actually allowed to add planes to the base format, as implied by the comment. Signed-off-by: Daniel Stone <daniels@collabora.com> Reviewed-by: Tapani Pälli <tapani.palli@intel.com> Reviewed-by: Philipp Zabel <p.zabel@pengutronix.de>	2017-08-11 10:25:53 +01:00
Iago Toral Quiroga	81615ad444	intel/compiler: properly size attribute wa_flags array for Vulkan Mesa will map user defined vertex input attributes to slots starting at VERT_ATTRIB_GENERIC0 which gives us room for only 16 slots (up to GL_VERT_ATTRIB_MAX). This sufficient for GL, where we expose exactly 16 vertex attributes for user defined inputs, but in Vulkan we can expose up to 28 (which are also mapped from VERT_ATTRIB_GENERIC0 onwards) so we need to account for this when we scope the size of the array of attribute workaround flags that is used during the brw_vertex_workarounds NIR pass. This prevents out-of-bounds accesses in that array for NIR shaders that use more than 16 vertex input attributes. Fixes: dEQP-VK.pipeline.vertex_input.max_attributes.* Acked-by: Lionel Landwerlin <lionel.g.landwerlin@intel.com>	2017-08-11 10:41:44 +02:00
Timothy Arceri	9d41ec2182	glsl: stop cloning builtin fuctions _mesa_glsl_find_builtin_function() The cloning was introduced in `f81ede4699` to fix a problem with shaders including IR that was owned by builtins. However the approach of cloning the whole function each time we reference a builtin lead to a significant reduction in the GLSL IR compilers performance. The previous patch fixes the ownership problem in a more precise way. So we can now remove this cloning. Testing on a Ryzen 7 1800X shows a ~15% decreases in compiling the Deus Ex: Mankind Divided shaders on radeonsi (which take 5min+ on some machines). Looking just at the GLSL IR compiler the speed up is ~40%. Tested-by: Dieter Nützel <Dieter@nuetzel-hh.de> Reviewed-by: Kenneth Graunke <kenneth@whitecape.org> Reviewed-by: Nicolai Hähnle <nicolai.haehnle@amd.com>	2017-08-11 15:44:15 +10:00
Timothy Arceri	77f5221233	glsl: pass mem_ctx to constant_expression_value(...) and friends The main motivation for this is that threaded compilation can fall over if we were to allocate IR inside constant_expression_value() when calling it on a builtin. This is because builtins are shared across the whole OpenGL context. `f81ede4699` worked around the problem by cloning the entire builtin before constant_expression_value() could be called on it. However cloning the whole function each time we referenced it lead to a significant reduction in the GLSL IR compiler performance. This change along with the following patch helps fix that performance regression. Other advantages are that we reduce the number of calls to ralloc_parent(), and for loop unrolling we free constants after they are used rather than leaving them hanging around. Reviewed-by: Kenneth Graunke <kenneth@whitecape.org>	2017-08-11 15:44:08 +10:00

1 2 3 4 5 ...

87366 Commits