LLVM 24.0.0git
AArch64FrameLowering.cpp
Go to the documentation of this file.
1//===- AArch64FrameLowering.cpp - AArch64 Frame Lowering -------*- C++ -*-====//
2//
3// Part of the LLVM Project, under the Apache License v2.0 with LLVM Exceptions.
4// See https://llvm.org/LICENSE.txt for license information.
5// SPDX-License-Identifier: Apache-2.0 WITH LLVM-exception
6//
7//===----------------------------------------------------------------------===//
8//
9// This file contains the AArch64 implementation of TargetFrameLowering class.
10//
11// On AArch64, stack frames are structured as follows:
12//
13// The stack grows downward.
14//
15// All of the individual frame areas on the frame below are optional, i.e. it's
16// possible to create a function so that the particular area isn't present
17// in the frame.
18//
19// At function entry, the "frame" looks as follows:
20//
21// | | Higher address
22// |-----------------------------------|
23// | |
24// | arguments passed on the stack |
25// | |
26// |-----------------------------------| <- sp
27// | | Lower address
28//
29//
30// After the prologue has run, the frame has the following general structure.
31// Note that this doesn't depict the case where a red-zone is used. Also,
32// technically the last frame area (VLAs) doesn't get created until in the
33// main function body, after the prologue is run. However, it's depicted here
34// for completeness.
35//
36// | | Higher address
37// |-----------------------------------|
38// | |
39// | arguments passed on the stack |
40// | |
41// |-----------------------------------|
42// | |
43// | (Win64 only) varargs from reg |
44// | |
45// |-----------------------------------|
46// | |
47// | (Win64 only) callee-saved SVE reg |
48// | |
49// |-----------------------------------|
50// | |
51// | callee-saved gpr registers | <--.
52// | | | On Darwin platforms these
53// |- - - - - - - - - - - - - - - - - -| | callee saves are swapped,
54// | prev_lr | | (frame record first)
55// | prev_fp | <--'
56// | async context if needed |
57// | (a.k.a. "frame record") |
58// |-----------------------------------| <- fp(=x29)
59// Default SVE stack layout Split SVE objects
60// (aarch64-split-sve-objects=false) (aarch64-split-sve-objects=true)
61// |-----------------------------------| |-----------------------------------|
62// | <hazard padding> | | callee-saved PPR registers |
63// |-----------------------------------| |-----------------------------------|
64// | | | PPR stack objects |
65// | callee-saved fp/simd/SVE regs | |-----------------------------------|
66// | | | <hazard padding> |
67// |-----------------------------------| |-----------------------------------|
68// | | | callee-saved ZPR/FPR registers |
69// | SVE stack objects | |-----------------------------------|
70// | | | ZPR stack objects |
71// |-----------------------------------| |-----------------------------------|
72// ^ NB: FPR CSRs are promoted to ZPRs
73// |-----------------------------------|
74// |.empty.space.to.make.part.below....|
75// |.aligned.in.case.it.needs.more.than| (size of this area is unknown at
76// |.the.standard.16-byte.alignment....| compile time; if present)
77// |-----------------------------------|
78// | local variables of fixed size |
79// | including spill slots |
80// | <FPR> |
81// | <hazard padding> |
82// | <GPR> |
83// |-----------------------------------| <- bp(not defined by ABI,
84// |.variable-sized.local.variables....| LLVM chooses X19)
85// |.(VLAs)............................| (size of this area is unknown at
86// |...................................| compile time)
87// |-----------------------------------| <- sp
88// | | Lower address
89//
90//
91// To access the data in a frame, at-compile time, a constant offset must be
92// computable from one of the pointers (fp, bp, sp) to access it. The size
93// of the areas with a dotted background cannot be computed at compile-time
94// if they are present, making it required to have all three of fp, bp and
95// sp to be set up to be able to access all contents in the frame areas,
96// assuming all of the frame areas are non-empty.
97//
98// For most functions, some of the frame areas are empty. For those functions,
99// it may not be necessary to set up fp or bp:
100// * A base pointer is definitely needed when there are both VLAs and local
101// variables with more-than-default alignment requirements.
102// * A frame pointer is definitely needed when there are local variables with
103// more-than-default alignment requirements.
104//
105// For Darwin platforms the frame-record (fp, lr) is stored at the top of the
106// callee-saved area, since the unwind encoding does not allow for encoding
107// this dynamically and existing tools depend on this layout. For other
108// platforms, the frame-record is stored at the bottom of the (gpr) callee-saved
109// area to allow SVE stack objects (allocated directly below the callee-saves,
110// if available) to be accessed directly from the framepointer.
111// The SVE spill/fill instructions have VL-scaled addressing modes such
112// as:
113// ldr z8, [fp, #-7 mul vl]
114// For SVE the size of the vector length (VL) is not known at compile-time, so
115// '#-7 mul vl' is an offset that can only be evaluated at runtime. With this
116// layout, we don't need to add an unscaled offset to the framepointer before
117// accessing the SVE object in the frame.
118//
119// In some cases when a base pointer is not strictly needed, it is generated
120// anyway when offsets from the frame pointer to access local variables become
121// so large that the offset can't be encoded in the immediate fields of loads
122// or stores.
123//
124// Outgoing function arguments must be at the bottom of the stack frame when
125// calling another function. If we do not have variable-sized stack objects, we
126// can allocate a "reserved call frame" area at the bottom of the local
127// variable area, large enough for all outgoing calls. If we do have VLAs, then
128// the stack pointer must be decremented and incremented around each call to
129// make space for the arguments below the VLAs.
130//
131// FIXME: also explain the redzone concept.
132//
133// About stack hazards: Under some SME contexts, a coprocessor with its own
134// separate cache can used for FP operations. This can create hazards if the CPU
135// and the SME unit try to access the same area of memory, including if the
136// access is to an area of the stack. To try to alleviate this we attempt to
137// introduce extra padding into the stack frame between FP and GPR accesses,
138// controlled by the aarch64-stack-hazard-size option. Without changing the
139// layout of the stack frame in the diagram above, a stack object of size
140// aarch64-stack-hazard-size is added between GPR and FPR CSRs. Another is added
141// to the stack objects section, and stack objects are sorted so that FPR >
142// Hazard padding slot > GPRs (where possible). Unfortunately some things are
143// not handled well (VLA area, arguments on the stack, objects with both GPR and
144// FPR accesses), but if those are controlled by the user then the entire stack
145// frame becomes GPR at the start/end with FPR in the middle, surrounded by
146// Hazard padding.
147//
148// An example of the prologue:
149//
150// .globl __foo
151// .align 2
152// __foo:
153// Ltmp0:
154// .cfi_startproc
155// .cfi_personality 155, ___gxx_personality_v0
156// Leh_func_begin:
157// .cfi_lsda 16, Lexception33
158//
159// stp xa,bx, [sp, -#offset]!
160// ...
161// stp x28, x27, [sp, #offset-32]
162// stp fp, lr, [sp, #offset-16]
163// add fp, sp, #offset - 16
164// sub sp, sp, #1360
165//
166// The Stack:
167// +-------------------------------------------+
168// 10000 | ........ | ........ | ........ | ........ |
169// 10004 | ........ | ........ | ........ | ........ |
170// +-------------------------------------------+
171// 10008 | ........ | ........ | ........ | ........ |
172// 1000c | ........ | ........ | ........ | ........ |
173// +===========================================+
174// 10010 | X28 Register |
175// 10014 | X28 Register |
176// +-------------------------------------------+
177// 10018 | X27 Register |
178// 1001c | X27 Register |
179// +===========================================+
180// 10020 | Frame Pointer |
181// 10024 | Frame Pointer |
182// +-------------------------------------------+
183// 10028 | Link Register |
184// 1002c | Link Register |
185// +===========================================+
186// 10030 | ........ | ........ | ........ | ........ |
187// 10034 | ........ | ........ | ........ | ........ |
188// +-------------------------------------------+
189// 10038 | ........ | ........ | ........ | ........ |
190// 1003c | ........ | ........ | ........ | ........ |
191// +-------------------------------------------+
192//
193// [sp] = 10030 :: >>initial value<<
194// sp = 10020 :: stp fp, lr, [sp, #-16]!
195// fp = sp == 10020 :: mov fp, sp
196// [sp] == 10020 :: stp x28, x27, [sp, #-16]!
197// sp == 10010 :: >>final value<<
198//
199// The frame pointer (w29) points to address 10020. If we use an offset of
200// '16' from 'w29', we get the CFI offsets of -8 for w30, -16 for w29, -24
201// for w27, and -32 for w28:
202//
203// Ltmp1:
204// .cfi_def_cfa w29, 16
205// Ltmp2:
206// .cfi_offset w30, -8
207// Ltmp3:
208// .cfi_offset w29, -16
209// Ltmp4:
210// .cfi_offset w27, -24
211// Ltmp5:
212// .cfi_offset w28, -32
213//
214//===----------------------------------------------------------------------===//
215
216#include "AArch64FrameLowering.h"
217#include "AArch64InstrInfo.h"
220#include "AArch64RegisterInfo.h"
221#include "AArch64SMEAttributes.h"
222#include "AArch64Subtarget.h"
225#include "llvm/ADT/ScopeExit.h"
226#include "llvm/ADT/SmallVector.h"
244#include "llvm/IR/Attributes.h"
245#include "llvm/IR/CallingConv.h"
246#include "llvm/IR/DataLayout.h"
247#include "llvm/IR/DebugLoc.h"
248#include "llvm/IR/Function.h"
249#include "llvm/MC/MCAsmInfo.h"
250#include "llvm/MC/MCDwarf.h"
252#include "llvm/Support/Debug.h"
258#include <cassert>
259#include <cstdint>
260#include <iterator>
261#include <optional>
262#include <vector>
263
264using namespace llvm;
265
266#define DEBUG_TYPE "frame-info"
267
268static cl::opt<bool> EnableRedZone("aarch64-redzone",
269 cl::desc("enable use of redzone on AArch64"),
270 cl::init(false), cl::Hidden);
271
273 "stack-tagging-merge-settag",
274 cl::desc("merge settag instruction in function epilog"), cl::init(true),
275 cl::Hidden);
276
277static cl::opt<bool> OrderFrameObjects("aarch64-order-frame-objects",
278 cl::desc("sort stack allocations"),
279 cl::init(true), cl::Hidden);
280
281static cl::opt<bool>
282 SplitSVEObjects("aarch64-split-sve-objects",
283 cl::desc("Split allocation of ZPR & PPR objects"),
284 cl::init(true), cl::Hidden);
285
287 "homogeneous-prolog-epilog", cl::Hidden,
288 cl::desc("Emit homogeneous prologue and epilogue for the size "
289 "optimization (default = off)"));
290
291// Stack hazard size for analysis remarks. StackHazardSize takes precedence.
293 StackHazardRemarkSize("aarch64-stack-hazard-remark-size", cl::init(0),
294 cl::Hidden);
295// Whether to insert padding into non-streaming functions (for testing).
296static cl::opt<bool>
297 StackHazardInNonStreaming("aarch64-stack-hazard-in-non-streaming",
298 cl::init(false), cl::Hidden);
299
301 "aarch64-disable-multivector-spill-fill",
302 cl::desc("Disable use of LD/ST pairs for SME2 or SVE2p1"), cl::init(false),
303 cl::Hidden);
304
305int64_t
307 MachineBasicBlock &MBB) const {
308 MachineBasicBlock::iterator MBBI = MBB.getLastNonDebugInstr();
310 bool IsTailCallReturn = (MBB.end() != MBBI)
312 : false;
313
314 int64_t ArgumentPopSize = 0;
315 if (IsTailCallReturn) {
316 MachineOperand &StackAdjust = MBBI->getOperand(1);
317
318 // For a tail-call in a callee-pops-arguments environment, some or all of
319 // the stack may actually be in use for the call's arguments, this is
320 // calculated during LowerCall and consumed here...
321 ArgumentPopSize = StackAdjust.getImm();
322 } else {
323 // ... otherwise the amount to pop is *all* of the argument space,
324 // conveniently stored in the MachineFunctionInfo by
325 // LowerFormalArguments. This will, of course, be zero for the C calling
326 // convention.
327 ArgumentPopSize = AFI->getArgumentStackToRestore();
328 }
329
330 return ArgumentPopSize;
331}
332
334 MachineFunction &MF);
335
336enum class AssignObjectOffsets { No, Yes };
337/// Process all the SVE stack objects and the SVE stack size and offsets for
338/// each object. If AssignOffsets is "Yes", the offsets get assigned (and SVE
339/// stack sizes set). Returns the size of the SVE stack.
341 AssignObjectOffsets AssignOffsets);
342
343static unsigned getStackHazardSize(const MachineFunction &MF) {
344 return MF.getSubtarget<AArch64Subtarget>().getStreamingHazardSize();
345}
346
352
355 // With split SVE objects, the hazard padding is added to the PPR region,
356 // which places it between the [GPR, PPR] area and the [ZPR, FPR] area. This
357 // avoids hazards between both GPRs and FPRs and ZPRs and PPRs.
360 : 0,
361 AFI->getStackSizePPR());
362}
363
364// Conservatively, returns true if the function is likely to have SVE vectors
365// on the stack. This function is safe to be called before callee-saves or
366// object offsets have been determined.
368 const MachineFunction &MF) {
369 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
370 if (AFI->isSVECC())
371 return true;
372
373 if (AFI->hasCalculatedStackSizeSVE())
374 return bool(AFL.getSVEStackSize(MF));
375
376 const MachineFrameInfo &MFI = MF.getFrameInfo();
377 for (int FI = MFI.getObjectIndexBegin(); FI < MFI.getObjectIndexEnd(); FI++) {
378 if (MFI.hasScalableStackID(FI))
379 return true;
380 }
381
382 return false;
383}
384
385static bool isTargetWindows(const MachineFunction &MF) {
386 // TODO: Should this include targets like UEFI (which use Windows CFI)?
387 // Note: Currently, there is not AArch64 support for UEFI. The value returned
388 // here must align with the predicate used for returning the list of callee
389 // saved regs in AArch64RegisterInfo::getCalleeSavedRegs(), so that we use
390 // invalidateWindowsRegisterPairing() where appropriate.
392}
393
395 const MachineFunction &MF) const {
396 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
397 return isTargetWindows(MF) && AFI->getSVECalleeSavedStackSize();
398}
399
400/// Returns true if a homogeneous prolog or epilog code can be emitted
401/// for the size optimization. If possible, a frame helper call is injected.
402/// When Exit block is given, this check is for epilog.
403bool AArch64FrameLowering::homogeneousPrologEpilog(
404 MachineFunction &MF, MachineBasicBlock *Exit) const {
405 if (!MF.getFunction().hasMinSize())
406 return false;
408 return false;
409 if (EnableRedZone)
410 return false;
411
412 // TODO: Window is supported yet.
413 if (isTargetWindows(MF))
414 return false;
415
416 // TODO: SVE is not supported yet.
417 if (isLikelyToHaveSVEStack(*this, MF))
418 return false;
419
420 // Bail on stack adjustment needed on return for simplicity.
421 const MachineFrameInfo &MFI = MF.getFrameInfo();
422 const TargetRegisterInfo *RegInfo = MF.getSubtarget().getRegisterInfo();
423 if (MFI.hasVarSizedObjects() || RegInfo->hasStackRealignment(MF))
424 return false;
425 if (Exit && getArgumentStackToRestore(MF, *Exit))
426 return false;
427
428 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
429 if (AFI->hasSwiftAsyncContext() || AFI->hasStreamingModeChanges())
430 return false;
431
432 // If there are an odd number of GPRs before LR and FP in the CSRs list,
433 // they will not be paired into one RegPairInfo, which is incompatible with
434 // the assumption made by the homogeneous prolog epilog pass.
435 const MCPhysReg *CSRegs = MF.getRegInfo().getCalleeSavedRegs();
436 unsigned NumGPRs = 0;
437 for (unsigned I = 0; CSRegs[I]; ++I) {
438 Register Reg = CSRegs[I];
439 if (Reg == AArch64::LR) {
440 assert(CSRegs[I + 1] == AArch64::FP);
441 if (NumGPRs % 2 != 0)
442 return false;
443 break;
444 }
445 if (AArch64::GPR64RegClass.contains(Reg))
446 ++NumGPRs;
447 }
448
449 return true;
450}
451
452/// Returns true if CSRs should be paired.
453bool AArch64FrameLowering::producePairRegisters(MachineFunction &MF) const {
454 return produceCompactUnwindFrame(*this, MF) || homogeneousPrologEpilog(MF);
455}
456
457/// This is the biggest offset to the stack pointer we can encode in aarch64
458/// instructions (without using a separate calculation and a temp register).
459/// Note that the exception here are vector stores/loads which cannot encode any
460/// displacements (see estimateRSStackSizeLimit(), isAArch64FrameOffsetLegal()).
461static const unsigned DefaultSafeSPDisplacement = 255;
462
463/// Look at each instruction that references stack frames and return the stack
464/// size limit beyond which some of these instructions will require a scratch
465/// register during their expansion later.
467 // FIXME: For now, just conservatively guesstimate based on unscaled indexing
468 // range. We'll end up allocating an unnecessary spill slot a lot, but
469 // realistically that's not a big deal at this stage of the game.
470 for (MachineBasicBlock &MBB : MF) {
471 for (MachineInstr &MI : MBB) {
472 if (MI.isDebugInstr() || MI.isPseudo() ||
473 MI.getOpcode() == AArch64::ADDXri ||
474 MI.getOpcode() == AArch64::ADDSXri)
475 continue;
476
477 for (const MachineOperand &MO : MI.operands()) {
478 if (!MO.isFI())
479 continue;
480
482 if (isAArch64FrameOffsetLegal(MI, Offset, nullptr, nullptr, nullptr) ==
484 return 0;
485 }
486 }
487 }
489}
490
495
496unsigned
497AArch64FrameLowering::getFixedObjectSize(const MachineFunction &MF,
498 const AArch64FunctionInfo *AFI,
499 bool IsWin64, bool IsFunclet) const {
500 assert(AFI->getTailCallReservedStack() % 16 == 0 &&
501 "Tail call reserved stack must be aligned to 16 bytes");
502 if (!IsWin64 || IsFunclet) {
503 return AFI->getTailCallReservedStack();
504 } else {
505 if (AFI->getTailCallReservedStack() != 0 &&
506 !MF.getFunction().getAttributes().hasAttrSomewhere(
507 Attribute::SwiftAsync))
508 report_fatal_error("cannot generate ABI-changing tail call for Win64");
509 unsigned FixedObjectSize = AFI->getTailCallReservedStack();
510
511 // Var args are stored here in the primary function.
512 FixedObjectSize += AFI->getVarArgsGPRSize();
513
514 if (MF.hasEHFunclets()) {
515 // Catch objects are stored here in the primary function.
516 const MachineFrameInfo &MFI = MF.getFrameInfo();
517 const WinEHFuncInfo &EHInfo = *MF.getWinEHFuncInfo();
518 SmallSetVector<int, 8> CatchObjFrameIndices;
519 for (const WinEHTryBlockMapEntry &TBME : EHInfo.TryBlockMap) {
520 for (const WinEHHandlerType &H : TBME.HandlerArray) {
521 int FrameIndex = H.CatchObj.FrameIndex;
522 if ((FrameIndex != INT_MAX) &&
523 CatchObjFrameIndices.insert(FrameIndex)) {
524 FixedObjectSize = alignTo(FixedObjectSize,
525 MFI.getObjectAlign(FrameIndex).value()) +
526 MFI.getObjectSize(FrameIndex);
527 }
528 }
529 }
530 // To support EH funclets we allocate an UnwindHelp object
531 FixedObjectSize += 8;
532 }
533 return alignTo(FixedObjectSize, 16);
534 }
535}
536
538 if (!EnableRedZone)
539 return false;
540
541 // Don't use the red zone if the function explicitly asks us not to.
542 // This is typically used for kernel code.
543 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
544 const unsigned RedZoneSize =
546 if (!RedZoneSize)
547 return false;
548
549 const MachineFrameInfo &MFI = MF.getFrameInfo();
551 uint64_t NumBytes = AFI->getLocalStackSize();
552
553 // If neither NEON or SVE are available, a COPY from one Q-reg to
554 // another requires a spill -> reload sequence. We can do that
555 // using a pre-decrementing store/post-decrementing load, but
556 // if we do so, we can't use the Red Zone.
557 bool LowerQRegCopyThroughMem = Subtarget.hasFPARMv8() &&
558 !Subtarget.isNeonAvailable() &&
559 !Subtarget.hasSVE();
560
561 return !(MFI.hasCalls() || hasFP(MF) || NumBytes > RedZoneSize ||
562 AFI->hasSVEStackSize() || LowerQRegCopyThroughMem);
563}
564
565/// hasFPImpl - Return true if the specified function should have a dedicated
566/// frame pointer register.
568 const MachineFrameInfo &MFI = MF.getFrameInfo();
569 const TargetRegisterInfo *RegInfo = MF.getSubtarget().getRegisterInfo();
571
572 // Win64 EH requires a frame pointer if funclets are present, as the locals
573 // are accessed off the frame pointer in both the parent function and the
574 // funclets.
575 if (MF.hasEHFunclets())
576 return true;
577
578 // When the stack guard is mixed with the frame pointer, a dedicated FP is
579 // required so the guard value remains stable in the presence of dynamic
580 // stack allocations (e.g. _alloca on MSVCRT).
581 if (MFI.hasStackProtectorIndex()) {
582 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
583 if (Subtarget.getTargetLowering()->useStackGuardMixFP())
584 return true;
585 }
586
587 // Retain behavior of always omitting the FP for leaf functions when possible.
589 return true;
590 if (MFI.hasVarSizedObjects() || MFI.isFrameAddressTaken() ||
591 MFI.hasStackMap() || MFI.hasPatchPoint() ||
592 RegInfo->hasStackRealignment(MF))
593 return true;
594
595 // If we:
596 //
597 // 1. Have streaming mode changes
598 // OR:
599 // 2. Have a streaming body with SVE stack objects
600 //
601 // Then the value of VG restored when unwinding to this function may not match
602 // the value of VG used to set up the stack.
603 //
604 // This is a problem as the CFA can be described with an expression of the
605 // form: CFA = SP + NumBytes + VG * NumScalableBytes.
606 //
607 // If the value of VG used in that expression does not match the value used to
608 // set up the stack, an incorrect address for the CFA will be computed, and
609 // unwinding will fail.
610 //
611 // We work around this issue by ensuring the frame-pointer can describe the
612 // CFA in either of these cases.
613 if (AFI.needsDwarfUnwindInfo(MF) &&
616 return true;
617 // With large callframes around we may need to use FP to access the scavenging
618 // emergency spillslot.
619 //
620 // Unfortunately some calls to hasFP() like machine verifier ->
621 // getReservedReg() -> hasFP in the middle of global isel are too early
622 // to know the max call frame size. Hopefully conservatively returning "true"
623 // in those cases is fine.
624 // DefaultSafeSPDisplacement is fine as we only emergency spill GP regs.
625 if (!MFI.isMaxCallFrameSizeComputed() ||
627 return true;
628
629 return false;
630}
631
632/// Should the Frame Pointer be reserved for the current function?
634 const TargetMachine &TM = MF.getTarget();
635 const Triple &TT = TM.getTargetTriple();
636
637 // These OSes require the frame chain is valid, even if the current frame does
638 // not use a frame pointer.
639 if (TT.isOSDarwin() || TT.isOSWindows())
640 return true;
641
642 // If the function has a frame pointer, it is reserved.
643 if (hasFP(MF))
644 return true;
645
646 // Frontend has requested to preserve the frame pointer.
647 if (MF.framePointerIsReserved())
648 return true;
649
650 return false;
651}
652
653/// hasReservedCallFrame - Under normal circumstances, when a frame pointer is
654/// not required, we reserve argument space for call sites in the function
655/// immediately on entry to the current function. This eliminates the need for
656/// add/sub sp brackets around call sites. Returns true if the call frame is
657/// included as part of the stack frame.
659 const MachineFunction &MF) const {
660 // The stack probing code for the dynamically allocated outgoing arguments
661 // area assumes that the stack is probed at the top - either by the prologue
662 // code, which issues a probe if `hasVarSizedObjects` return true, or by the
663 // most recent variable-sized object allocation. Changing the condition here
664 // may need to be followed up by changes to the probe issuing logic.
665 return !MF.getFrameInfo().hasVarSizedObjects();
666}
667
671
672 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
673 const AArch64InstrInfo *TII = Subtarget.getInstrInfo();
674 const AArch64TargetLowering *TLI = Subtarget.getTargetLowering();
675 [[maybe_unused]] MachineFrameInfo &MFI = MF.getFrameInfo();
676 DebugLoc DL = I->getDebugLoc();
677 unsigned Opc = I->getOpcode();
678 bool IsDestroy = Opc == TII->getCallFrameDestroyOpcode();
679 uint64_t CalleePopAmount = IsDestroy ? I->getOperand(1).getImm() : 0;
680
681 if (!hasReservedCallFrame(MF)) {
682 int64_t Amount = I->getOperand(0).getImm();
683 Amount = alignTo(Amount, getStackAlign());
684 if (!IsDestroy)
685 Amount = -Amount;
686
687 // N.b. if CalleePopAmount is valid but zero (i.e. callee would pop, but it
688 // doesn't have to pop anything), then the first operand will be zero too so
689 // this adjustment is a no-op.
690 if (CalleePopAmount == 0) {
691 // FIXME: in-function stack adjustment for calls is limited to 24-bits
692 // because there's no guaranteed temporary register available.
693 //
694 // ADD/SUB (immediate) has only LSL #0 and LSL #12 available.
695 // 1) For offset <= 12-bit, we use LSL #0
696 // 2) For 12-bit <= offset <= 24-bit, we use two instructions. One uses
697 // LSL #0, and the other uses LSL #12.
698 //
699 // Most call frames will be allocated at the start of a function so
700 // this is OK, but it is a limitation that needs dealing with.
701 assert(Amount > -0xffffff && Amount < 0xffffff && "call frame too large");
702
703 if (TLI->hasInlineStackProbe(MF) &&
705 // When stack probing is enabled, the decrement of SP may need to be
706 // probed. We only need to do this if the call site needs 1024 bytes of
707 // space or more, because a region smaller than that is allowed to be
708 // unprobed at an ABI boundary. We rely on the fact that SP has been
709 // probed exactly at this point, either by the prologue or most recent
710 // dynamic allocation.
712 "non-reserved call frame without var sized objects?");
713 Register ScratchReg =
714 MF.getRegInfo().createVirtualRegister(&AArch64::GPR64RegClass);
715 inlineStackProbeFixed(I, ScratchReg, -Amount, StackOffset::get(0, 0));
716 } else {
717 emitFrameOffset(MBB, I, DL, AArch64::SP, AArch64::SP,
718 StackOffset::getFixed(Amount), TII);
719 }
720 }
721 } else if (CalleePopAmount != 0) {
722 // If the calling convention demands that the callee pops arguments from the
723 // stack, we want to add it back if we have a reserved call frame.
724 assert(CalleePopAmount < 0xffffff && "call frame too large");
725 emitFrameOffset(MBB, I, DL, AArch64::SP, AArch64::SP,
726 StackOffset::getFixed(-(int64_t)CalleePopAmount), TII);
727 }
728 return MBB.erase(I);
729}
730
732 MachineBasicBlock &MBB) const {
733
734 MachineFunction &MF = *MBB.getParent();
735 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
736 const auto &TRI = *Subtarget.getRegisterInfo();
737 const auto &MFI = *MF.getInfo<AArch64FunctionInfo>();
738
739 CFIInstBuilder CFIBuilder(MBB, MBB.begin(), MachineInstr::NoFlags);
740
741 // Reset the CFA to `SP + 0`.
742 CFIBuilder.buildDefCFA(AArch64::SP, 0);
743
744 // Flip the RA sign state.
745 if (MFI.shouldSignReturnAddress(MF)) {
746 if (MFI.branchProtectionPAuthLR()) {
747 CFIBuilder.buildNegateRAStateWithPC();
748 } else if (!MF.getTarget().getTargetTriple().isOSBinFormatMachO()) {
749 CFIBuilder.buildNegateRAState();
750 }
751 }
752
753 // Shadow call stack uses X18, reset it.
754 if (MFI.needsShadowCallStackPrologueEpilogue(MF))
755 CFIBuilder.buildSameValue(AArch64::X18);
756
757 // Emit .cfi_same_value for callee-saved registers.
758 const std::vector<CalleeSavedInfo> &CSI =
760 for (const auto &Info : CSI) {
761 MCRegister Reg = Info.getReg();
762 if (!TRI.regNeedsCFI(Reg, Reg))
763 continue;
764 CFIBuilder.buildSameValue(Reg);
765 }
766}
767
769 switch (Reg.id()) {
770 default:
771 // The called routine is expected to preserve r19-r28
772 // r29 and r30 are used as frame pointer and link register resp.
773 return 0;
774
775 // GPRs
776#define CASE(n) \
777 case AArch64::W##n: \
778 case AArch64::X##n: \
779 return AArch64::X##n
780 CASE(0);
781 CASE(1);
782 CASE(2);
783 CASE(3);
784 CASE(4);
785 CASE(5);
786 CASE(6);
787 CASE(7);
788 CASE(8);
789 CASE(9);
790 CASE(10);
791 CASE(11);
792 CASE(12);
793 CASE(13);
794 CASE(14);
795 CASE(15);
796 CASE(16);
797 CASE(17);
798 CASE(18);
799#undef CASE
800
801 // FPRs
802#define CASE(n) \
803 case AArch64::B##n: \
804 case AArch64::H##n: \
805 case AArch64::S##n: \
806 case AArch64::D##n: \
807 case AArch64::Q##n: \
808 return HasSVE ? AArch64::Z##n : AArch64::Q##n
809 CASE(0);
810 CASE(1);
811 CASE(2);
812 CASE(3);
813 CASE(4);
814 CASE(5);
815 CASE(6);
816 CASE(7);
817 CASE(8);
818 CASE(9);
819 CASE(10);
820 CASE(11);
821 CASE(12);
822 CASE(13);
823 CASE(14);
824 CASE(15);
825 CASE(16);
826 CASE(17);
827 CASE(18);
828 CASE(19);
829 CASE(20);
830 CASE(21);
831 CASE(22);
832 CASE(23);
833 CASE(24);
834 CASE(25);
835 CASE(26);
836 CASE(27);
837 CASE(28);
838 CASE(29);
839 CASE(30);
840 CASE(31);
841#undef CASE
842 }
843}
844
845void AArch64FrameLowering::emitZeroCallUsedRegs(BitVector RegsToZero,
847 RegScavenger *) const {
848 // Insertion point.
850
851 // Fake a debug loc.
852 DebugLoc DL;
853 if (MBBI != MBB.end())
854 DL = MBBI->getDebugLoc();
855
856 const MachineFunction &MF = *MBB.getParent();
857 const AArch64Subtarget &STI = MF.getSubtarget<AArch64Subtarget>();
858 const AArch64RegisterInfo &TRI = *STI.getRegisterInfo();
859
860 BitVector GPRsToZero(TRI.getNumRegs());
861 BitVector FPRsToZero(TRI.getNumRegs());
862 bool HasSVE = STI.isSVEorStreamingSVEAvailable();
863 // Without an FP unit (e.g. -mgeneral-regs-only) the FP/vector registers can't
864 // hold a value and there is no instruction to clear them, so leave them out.
865 bool HasFPR = STI.hasFPARMv8();
866 for (MCRegister Reg : RegsToZero.set_bits()) {
867 if (TRI.isGeneralPurposeRegister(MF, Reg)) {
868 // For GPRs, we only care to clear out the 64-bit register.
869 if (MCRegister XReg = getRegisterOrZero(Reg, HasSVE))
870 GPRsToZero.set(XReg);
871 } else if (HasFPR && AArch64InstrInfo::isFpOrNEON(Reg)) {
872 // For FPRs,
873 if (MCRegister XReg = getRegisterOrZero(Reg, HasSVE))
874 FPRsToZero.set(XReg);
875 }
876 }
877
878 const AArch64InstrInfo &TII = *STI.getInstrInfo();
879
880 // Zero out GPRs.
881 for (MCRegister Reg : GPRsToZero.set_bits())
882 TII.buildClearRegister(Reg, MBB, MBBI, DL);
883
884 // Zero out FP/vector registers.
885 for (MCRegister Reg : FPRsToZero.set_bits())
886 TII.buildClearRegister(Reg, MBB, MBBI, DL);
887
888 if (HasSVE) {
889 for (MCRegister PReg :
890 {AArch64::P0, AArch64::P1, AArch64::P2, AArch64::P3, AArch64::P4,
891 AArch64::P5, AArch64::P6, AArch64::P7, AArch64::P8, AArch64::P9,
892 AArch64::P10, AArch64::P11, AArch64::P12, AArch64::P13, AArch64::P14,
893 AArch64::P15}) {
894 if (RegsToZero[PReg])
895 BuildMI(MBB, MBBI, DL, TII.get(AArch64::PFALSE), PReg);
896 }
897 }
898}
899
900bool AArch64FrameLowering::windowsRequiresStackProbe(
901 const MachineFunction &MF, uint64_t StackSizeInBytes) const {
902 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
903 const AArch64FunctionInfo &MFI = *MF.getInfo<AArch64FunctionInfo>();
904 // TODO: When implementing stack protectors, take that into account
905 // for the probe threshold.
906 return Subtarget.isTargetWindows() && MFI.hasStackProbing() &&
907 StackSizeInBytes >= uint64_t(MFI.getStackProbeSize());
908}
909
911 const MachineBasicBlock &MBB) {
912 const MachineFunction *MF = MBB.getParent();
913 LiveRegs.addLiveIns(MBB);
914 // Mark callee saved registers as used so we will not choose them.
915 const MCPhysReg *CSRegs = MF->getRegInfo().getCalleeSavedRegs();
916 for (unsigned i = 0; CSRegs[i]; ++i)
917 LiveRegs.addReg(CSRegs[i]);
918}
919
921AArch64FrameLowering::findScratchNonCalleeSaveRegister(MachineBasicBlock *MBB,
922 bool HasCall) const {
924
925 // If MBB is an entry block, use X9 as the scratch register
926 // preserve_none functions may be using X9 to pass arguments,
927 // so prefer to pick an available register below.
928 if (&MF->front() == MBB &&
930 return AArch64::X9;
931
932 const AArch64Subtarget &Subtarget = MF->getSubtarget<AArch64Subtarget>();
933 const AArch64RegisterInfo &TRI = *Subtarget.getRegisterInfo();
934 LivePhysRegs LiveRegs(TRI);
935 getLiveRegsForEntryMBB(LiveRegs, *MBB);
936 if (HasCall) {
937 LiveRegs.addReg(AArch64::X16);
938 LiveRegs.addReg(AArch64::X17);
939 LiveRegs.addReg(AArch64::X18);
940 }
941
942 // Prefer X9 since it was historically used for the prologue scratch reg.
943 const MachineRegisterInfo &MRI = MF->getRegInfo();
944 if (LiveRegs.available(MRI, AArch64::X9))
945 return AArch64::X9;
946
947 for (unsigned Reg : AArch64::GPR64RegClass) {
948 if (LiveRegs.available(MRI, Reg))
949 return Reg;
950 }
951 return AArch64::NoRegister;
952}
953
955 const MachineBasicBlock &MBB) const {
956 const MachineFunction *MF = MBB.getParent();
957 MachineBasicBlock *TmpMBB = const_cast<MachineBasicBlock *>(&MBB);
958 const AArch64Subtarget &Subtarget = MF->getSubtarget<AArch64Subtarget>();
959 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
960 const AArch64TargetLowering *TLI = Subtarget.getTargetLowering();
962
963 if (AFI->hasSwiftAsyncContext()) {
964 const AArch64RegisterInfo &TRI = *Subtarget.getRegisterInfo();
965 const MachineRegisterInfo &MRI = MF->getRegInfo();
968 // The StoreSwiftAsyncContext clobbers X16 and X17. Make sure they are
969 // available.
970 if (!LiveRegs.available(MRI, AArch64::X16) ||
971 !LiveRegs.available(MRI, AArch64::X17))
972 return false;
973 }
974
975 // Certain stack probing sequences might clobber flags, then we can't use
976 // the block as a prologue if the flags register is a live-in.
978 MBB.isLiveIn(AArch64::NZCV))
979 return false;
980
981 if (RegInfo->hasStackRealignment(*MF) || TLI->hasInlineStackProbe(*MF))
982 if (findScratchNonCalleeSaveRegister(TmpMBB) == AArch64::NoRegister)
983 return false;
984
985 // May need a scratch register (for return value) if require making a special
986 // call
987 if (requiresSaveVG(*MF) ||
988 windowsRequiresStackProbe(*MF, std::numeric_limits<uint64_t>::max()))
989 if (findScratchNonCalleeSaveRegister(TmpMBB, true) == AArch64::NoRegister)
990 return false;
991
992 return true;
993}
994
996 const Function &F = MF.getFunction();
997 return MF.getTarget().getMCAsmInfo().usesWindowsCFI() &&
998 F.needsUnwindTableEntry();
999}
1000
1001bool AArch64FrameLowering::shouldSignReturnAddressEverywhere(
1002 const MachineFunction &MF) const {
1003 // FIXME: With WinCFI, extra care should be taken to place SEH_PACSignLR
1004 // and SEH_EpilogEnd instructions in the correct order.
1006 return false;
1009}
1010
1011// Given a load or a store instruction, generate an appropriate unwinding SEH
1012// code on Windows.
1014AArch64FrameLowering::insertSEH(MachineBasicBlock::iterator MBBI,
1015 const AArch64InstrInfo &TII,
1016 MachineInstr::MIFlag Flag) const {
1017 unsigned Opc = MBBI->getOpcode();
1018 MachineBasicBlock *MBB = MBBI->getParent();
1019 MachineFunction &MF = *MBB->getParent();
1020 DebugLoc DL = MBBI->getDebugLoc();
1021 unsigned ImmIdx = MBBI->getNumOperands() - 1;
1022 int Imm = MBBI->getOperand(ImmIdx).getImm();
1024 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1025 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
1026
1027 switch (Opc) {
1028 default:
1029 report_fatal_error("No SEH Opcode for this instruction");
1030 case AArch64::STR_ZXI:
1031 case AArch64::LDR_ZXI: {
1032 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1033 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveZReg))
1034 .addImm(Reg0)
1035 .addImm(Imm)
1036 .setMIFlag(Flag);
1037 break;
1038 }
1039 case AArch64::STR_PXI:
1040 case AArch64::LDR_PXI: {
1041 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1042 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SavePReg))
1043 .addImm(Reg0)
1044 .addImm(Imm)
1045 .setMIFlag(Flag);
1046 break;
1047 }
1048 case AArch64::LDPDpost:
1049 Imm = -Imm;
1050 [[fallthrough]];
1051 case AArch64::STPDpre: {
1052 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1053 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(2).getReg());
1054 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFRegP_X))
1055 .addImm(Reg0)
1056 .addImm(Reg1)
1057 .addImm(Imm * 8)
1058 .setMIFlag(Flag);
1059 break;
1060 }
1061 case AArch64::LDPXpost:
1062 Imm = -Imm;
1063 [[fallthrough]];
1064 case AArch64::STPXpre: {
1065 Register Reg0 = MBBI->getOperand(1).getReg();
1066 Register Reg1 = MBBI->getOperand(2).getReg();
1067 if (Reg0 == AArch64::FP && Reg1 == AArch64::LR)
1068 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFPLR_X))
1069 .addImm(Imm * 8)
1070 .setMIFlag(Flag);
1071 else
1072 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveRegP_X))
1073 .addImm(RegInfo->getSEHRegNum(Reg0))
1074 .addImm(RegInfo->getSEHRegNum(Reg1))
1075 .addImm(Imm * 8)
1076 .setMIFlag(Flag);
1077 break;
1078 }
1079 case AArch64::LDRDpost:
1080 Imm = -Imm;
1081 [[fallthrough]];
1082 case AArch64::STRDpre: {
1083 unsigned Reg = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1084 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFReg_X))
1085 .addImm(Reg)
1086 .addImm(Imm)
1087 .setMIFlag(Flag);
1088 break;
1089 }
1090 case AArch64::LDRXpost:
1091 Imm = -Imm;
1092 [[fallthrough]];
1093 case AArch64::STRXpre: {
1094 unsigned Reg = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1095 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveReg_X))
1096 .addImm(Reg)
1097 .addImm(Imm)
1098 .setMIFlag(Flag);
1099 break;
1100 }
1101 case AArch64::STPDi:
1102 case AArch64::LDPDi: {
1103 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1104 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1105 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFRegP))
1106 .addImm(Reg0)
1107 .addImm(Reg1)
1108 .addImm(Imm * 8)
1109 .setMIFlag(Flag);
1110 break;
1111 }
1112 case AArch64::STPXi:
1113 case AArch64::LDPXi: {
1114 Register Reg0 = MBBI->getOperand(0).getReg();
1115 Register Reg1 = MBBI->getOperand(1).getReg();
1116
1117 int SEHReg0 = RegInfo->getSEHRegNum(Reg0);
1118 int SEHReg1 = RegInfo->getSEHRegNum(Reg1);
1119
1120 if (Reg0 == AArch64::FP && Reg1 == AArch64::LR)
1121 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFPLR))
1122 .addImm(Imm * 8)
1123 .setMIFlag(Flag);
1124 else if (SEHReg0 >= 19 && SEHReg1 >= 19)
1125 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveRegP))
1126 .addImm(SEHReg0)
1127 .addImm(SEHReg1)
1128 .addImm(Imm * 8)
1129 .setMIFlag(Flag);
1130 else
1131 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegIP))
1132 .addImm(SEHReg0)
1133 .addImm(SEHReg1)
1134 .addImm(Imm * 8)
1135 .setMIFlag(Flag);
1136 break;
1137 }
1138 case AArch64::STRXui:
1139 case AArch64::LDRXui: {
1140 int Reg = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1141 if (Reg >= 19)
1142 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveReg))
1143 .addImm(Reg)
1144 .addImm(Imm * 8)
1145 .setMIFlag(Flag);
1146 else
1147 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegI))
1148 .addImm(Reg)
1149 .addImm(Imm * 8)
1150 .setMIFlag(Flag);
1151 break;
1152 }
1153 case AArch64::STRDui:
1154 case AArch64::LDRDui: {
1155 unsigned Reg = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1156 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveFReg))
1157 .addImm(Reg)
1158 .addImm(Imm * 8)
1159 .setMIFlag(Flag);
1160 break;
1161 }
1162 case AArch64::STPQi:
1163 case AArch64::LDPQi: {
1164 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(0).getReg());
1165 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1166 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegQP))
1167 .addImm(Reg0)
1168 .addImm(Reg1)
1169 .addImm(Imm * 16)
1170 .setMIFlag(Flag);
1171 break;
1172 }
1173 case AArch64::LDPQpost:
1174 Imm = -Imm;
1175 [[fallthrough]];
1176 case AArch64::STPQpre: {
1177 unsigned Reg0 = RegInfo->getSEHRegNum(MBBI->getOperand(1).getReg());
1178 unsigned Reg1 = RegInfo->getSEHRegNum(MBBI->getOperand(2).getReg());
1179 MIB = BuildMI(MF, DL, TII.get(AArch64::SEH_SaveAnyRegQPX))
1180 .addImm(Reg0)
1181 .addImm(Reg1)
1182 .addImm(Imm * 16)
1183 .setMIFlag(Flag);
1184 break;
1185 }
1186 }
1187 auto I = MBB->insertAfter(MBBI, MIB);
1188 return I;
1189}
1190
1193 if (!AFI->needsDwarfUnwindInfo(MF) || !AFI->hasStreamingModeChanges())
1194 return false;
1195 // For Darwin platforms we don't save VG for non-SVE functions, even if SME
1196 // is enabled with streaming mode changes.
1197 auto &ST = MF.getSubtarget<AArch64Subtarget>();
1198 if (ST.isTargetDarwin())
1199 return ST.hasSVE();
1200 return true;
1201}
1202
1204 MachineFunction &MF) const {
1205 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1206 const AArch64InstrInfo *TII = Subtarget.getInstrInfo();
1207
1208 auto EmitSignRA = [&](MachineBasicBlock &MBB) {
1209 DebugLoc DL; // Set debug location to unknown.
1211
1212 BuildMI(MBB, MBBI, DL, TII->get(AArch64::PAUTH_PROLOGUE))
1214 };
1215
1216 auto EmitAuthRA = [&](MachineBasicBlock &MBB) {
1217 DebugLoc DL;
1218 MachineBasicBlock::iterator MBBI = MBB.getFirstTerminator();
1219 if (MBBI != MBB.end())
1220 DL = MBBI->getDebugLoc();
1221
1222 TII->createPauthEpilogueInstr(MBB, DL);
1223 };
1224
1225 // This should be in sync with PEIImpl::calculateSaveRestoreBlocks.
1226 EmitSignRA(MF.front());
1227 for (MachineBasicBlock &MBB : MF) {
1228 if (MBB.isEHFuncletEntry())
1229 EmitSignRA(MBB);
1230 if (MBB.isReturnBlock())
1231 EmitAuthRA(MBB);
1232 }
1233}
1234
1236 MachineBasicBlock &MBB) const {
1237 AArch64PrologueEmitter PrologueEmitter(MF, MBB, *this);
1238 PrologueEmitter.emitPrologue();
1239}
1240
1242 MachineBasicBlock &MBB) const {
1243 AArch64EpilogueEmitter EpilogueEmitter(MF, MBB, *this);
1244 EpilogueEmitter.emitEpilogue();
1245}
1246
1249 MF.getInfo<AArch64FunctionInfo>()->needsDwarfUnwindInfo(MF);
1250}
1251
1253 return enableCFIFixup(MF) &&
1254 MF.getInfo<AArch64FunctionInfo>()->needsAsyncDwarfUnwindInfo(MF);
1255}
1256
1257/// getFrameIndexReference - Provide a base+offset reference to an FI slot for
1258/// debug info. It's the same as what we use for resolving the code-gen
1259/// references for now. FIXME: This can go wrong when references are
1260/// SP-relative and simple call frames aren't used.
1263 Register &FrameReg) const {
1265 MF, FI, FrameReg,
1266 /*PreferFP=*/
1267 MF.getFunction().hasFnAttribute(Attribute::SanitizeHWAddress) ||
1268 MF.getFunction().hasFnAttribute(Attribute::SanitizeMemTag),
1269 /*ForSimm=*/false);
1270}
1271
1274 int FI) const {
1275 // This function serves to provide a comparable offset from a single reference
1276 // point (the value of SP at function entry) that can be used for analysis,
1277 // e.g. the stack-frame-layout analysis pass. It is not guaranteed to be
1278 // correct for all objects in the presence of VLA-area objects or dynamic
1279 // stack re-alignment.
1280
1281 const auto &MFI = MF.getFrameInfo();
1282
1283 int64_t ObjectOffset = MFI.getObjectOffset(FI);
1284 StackOffset ZPRStackSize = getZPRStackSize(MF);
1285 StackOffset PPRStackSize = getPPRStackSize(MF);
1286 StackOffset SVEStackSize = ZPRStackSize + PPRStackSize;
1287
1288 // For VLA-area objects, just emit an offset at the end of the stack frame.
1289 // Whilst not quite correct, these objects do live at the end of the frame and
1290 // so it is more useful for analysis for the offset to reflect this.
1291 if (MFI.isVariableSizedObjectIndex(FI)) {
1292 return StackOffset::getFixed(-((int64_t)MFI.getStackSize())) - SVEStackSize;
1293 }
1294
1295 // This is correct in the absence of any SVE stack objects.
1296 if (!SVEStackSize)
1297 return StackOffset::getFixed(ObjectOffset - getOffsetOfLocalArea());
1298
1299 const auto *AFI = MF.getInfo<AArch64FunctionInfo>();
1300 bool FPAfterSVECalleeSaves = hasSVECalleeSavesAboveFrameRecord(MF);
1301 if (MFI.hasScalableStackID(FI)) {
1302 if (FPAfterSVECalleeSaves &&
1303 -ObjectOffset <= (int64_t)AFI->getSVECalleeSavedStackSize()) {
1304 assert(!AFI->hasSplitSVEObjects() &&
1305 "split-sve-objects not supported with FPAfterSVECalleeSaves");
1306 return StackOffset::getScalable(ObjectOffset);
1307 }
1308 StackOffset AccessOffset{};
1309 // The scalable vectors are below (lower address) the scalable predicates
1310 // with split SVE objects, so we must subtract the size of the predicates.
1311 if (AFI->hasSplitSVEObjects() &&
1312 MFI.getStackID(FI) == TargetStackID::ScalableVector)
1313 AccessOffset = -PPRStackSize;
1314 return AccessOffset +
1315 StackOffset::get(-((int64_t)AFI->getCalleeSavedStackSize()),
1316 ObjectOffset);
1317 }
1318
1319 bool IsFixed = MFI.isFixedObjectIndex(FI);
1320 bool IsCSR =
1321 !IsFixed && ObjectOffset >= -((int)AFI->getCalleeSavedStackSize(MFI));
1322
1323 StackOffset ScalableOffset = {};
1324 if (!IsFixed && !IsCSR) {
1325 ScalableOffset = -SVEStackSize;
1326 } else if (FPAfterSVECalleeSaves && IsCSR) {
1327 ScalableOffset =
1329 }
1330
1331 return StackOffset::getFixed(ObjectOffset) + ScalableOffset;
1332}
1333
1339
1340StackOffset AArch64FrameLowering::getFPOffset(const MachineFunction &MF,
1341 int64_t ObjectOffset) const {
1342 const auto *AFI = MF.getInfo<AArch64FunctionInfo>();
1343 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1344 const Function &F = MF.getFunction();
1345 bool IsWin64 = Subtarget.isCallingConvWin64(F.getCallingConv(), F.isVarArg());
1346 unsigned FixedObject =
1347 getFixedObjectSize(MF, AFI, IsWin64, /*IsFunclet=*/false);
1348 int64_t CalleeSaveSize = AFI->getCalleeSavedStackSize(MF.getFrameInfo());
1349 int64_t FPAdjust =
1350 CalleeSaveSize - AFI->getCalleeSaveBaseToFrameRecordOffset();
1351 return StackOffset::getFixed(ObjectOffset + FixedObject + FPAdjust);
1352}
1353
1354StackOffset AArch64FrameLowering::getStackOffset(const MachineFunction &MF,
1355 int64_t ObjectOffset) const {
1356 const auto &MFI = MF.getFrameInfo();
1357 return StackOffset::getFixed(ObjectOffset + (int64_t)MFI.getStackSize());
1358}
1359
1360// TODO: This function currently does not work for scalable vectors.
1362 int FI) const {
1363 const AArch64RegisterInfo *RegInfo =
1364 MF.getSubtarget<AArch64Subtarget>().getRegisterInfo();
1365 int ObjectOffset = MF.getFrameInfo().getObjectOffset(FI);
1366 return RegInfo->getLocalAddressRegister(MF) == AArch64::FP
1367 ? getFPOffset(MF, ObjectOffset).getFixed()
1368 : getStackOffset(MF, ObjectOffset).getFixed();
1369}
1370
1372 const MachineFunction &MF, int FI, Register &FrameReg, bool PreferFP,
1373 bool ForSimm) const {
1374 const auto &MFI = MF.getFrameInfo();
1375 int64_t ObjectOffset = MFI.getObjectOffset(FI);
1376 bool isFixed = MFI.isFixedObjectIndex(FI);
1377 auto StackID = static_cast<TargetStackID::Value>(MFI.getStackID(FI));
1378 return resolveFrameOffsetReference(MF, ObjectOffset, isFixed, StackID,
1379 FrameReg, PreferFP, ForSimm);
1380}
1381
1383 const MachineFunction &MF, int64_t ObjectOffset, bool isFixed,
1384 TargetStackID::Value StackID, Register &FrameReg, bool PreferFP,
1385 bool ForSimm) const {
1386 const auto &MFI = MF.getFrameInfo();
1387 const auto &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1388 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
1389 const auto *AFI = MF.getInfo<AArch64FunctionInfo>();
1390
1391 int64_t FPOffset = getFPOffset(MF, ObjectOffset).getFixed();
1392 int64_t Offset = getStackOffset(MF, ObjectOffset).getFixed();
1393 bool isCSR =
1394 !isFixed && ObjectOffset >= -((int)AFI->getCalleeSavedStackSize(MFI));
1395 bool isSVE = MFI.isScalableStackID(StackID);
1396
1397 StackOffset ZPRStackSize = getZPRStackSize(MF);
1398 StackOffset PPRStackSize = getPPRStackSize(MF);
1399 StackOffset SVEStackSize = ZPRStackSize + PPRStackSize;
1400
1401 // Use frame pointer to reference fixed objects. Use it for locals if
1402 // there are VLAs or a dynamically realigned SP (and thus the SP isn't
1403 // reliable as a base). Make sure useFPForScavengingIndex() does the
1404 // right thing for the emergency spill slot.
1405 bool UseFP = false;
1406 if (AFI->hasStackFrame() && !isSVE) {
1407 // We shouldn't prefer using the FP to access fixed-sized stack objects when
1408 // there are scalable (SVE) objects in between the FP and the fixed-sized
1409 // objects.
1410 PreferFP &= !SVEStackSize;
1411
1412 // Note: Keeping the following as multiple 'if' statements rather than
1413 // merging to a single expression for readability.
1414 //
1415 // Argument access should always use the FP.
1416 if (isFixed) {
1417 UseFP = hasFP(MF);
1418 } else if (isCSR && RegInfo->hasStackRealignment(MF)) {
1419 // References to the CSR area must use FP if we're re-aligning the stack
1420 // since the dynamically-sized alignment padding is between the SP/BP and
1421 // the CSR area.
1422 assert(hasFP(MF) && "Re-aligned stack must have frame pointer");
1423 UseFP = true;
1424 } else if (hasFP(MF) && !RegInfo->hasStackRealignment(MF)) {
1425 // If the FPOffset is negative and we're producing a signed immediate, we
1426 // have to keep in mind that the available offset range for negative
1427 // offsets is smaller than for positive ones. If an offset is available
1428 // via the FP and the SP, use whichever is closest.
1429 bool FPOffsetFits = !ForSimm || FPOffset >= -256;
1430 PreferFP |= Offset > -FPOffset && !SVEStackSize;
1431
1432 if (FPOffset >= 0) {
1433 // If the FPOffset is positive, that'll always be best, as the SP/BP
1434 // will be even further away.
1435 UseFP = true;
1436 } else if (MFI.hasVarSizedObjects()) {
1437 // If we have variable sized objects, we can use either FP or BP, as the
1438 // SP offset is unknown. We can use the base pointer if we have one and
1439 // FP is not preferred. If not, we're stuck with using FP.
1440 bool CanUseBP = RegInfo->hasBasePointer(MF);
1441 if (FPOffsetFits && CanUseBP) // Both are ok. Pick the best.
1442 UseFP = PreferFP;
1443 else if (!CanUseBP) // Can't use BP. Forced to use FP.
1444 UseFP = true;
1445 // else we can use BP and FP, but the offset from FP won't fit.
1446 // That will make us scavenge registers which we can probably avoid by
1447 // using BP. If it won't fit for BP either, we'll scavenge anyway.
1448 } else if (MF.hasEHFunclets() && !RegInfo->hasBasePointer(MF)) {
1449 // Funclets access the locals contained in the parent's stack frame
1450 // via the frame pointer, so we have to use the FP in the parent
1451 // function.
1452 (void) Subtarget;
1453 assert(Subtarget.isCallingConvWin64(MF.getFunction().getCallingConv(),
1454 MF.getFunction().isVarArg()) &&
1455 "Funclets should only be present on Win64");
1456 UseFP = true;
1457 } else {
1458 // We have the choice between FP and (SP or BP).
1459 if (FPOffsetFits && PreferFP) // If FP is the best fit, use it.
1460 UseFP = true;
1461 }
1462 }
1463 }
1464
1465 assert(
1466 ((isFixed || isCSR) || !RegInfo->hasStackRealignment(MF) || !UseFP) &&
1467 "In the presence of dynamic stack pointer realignment, "
1468 "non-argument/CSR objects cannot be accessed through the frame pointer");
1469
1470 bool FPAfterSVECalleeSaves = hasSVECalleeSavesAboveFrameRecord(MF);
1471
1472 if (isSVE) {
1473 StackOffset FPOffset = StackOffset::get(
1474 -AFI->getCalleeSaveBaseToFrameRecordOffset(), ObjectOffset);
1475 StackOffset SPOffset =
1476 SVEStackSize +
1477 StackOffset::get(MFI.getStackSize() - AFI->getCalleeSavedStackSize(),
1478 ObjectOffset);
1479
1480 // With split SVE objects the ObjectOffset is relative to the split area
1481 // (i.e. the PPR area or ZPR area respectively).
1482 if (AFI->hasSplitSVEObjects() && StackID == TargetStackID::ScalableVector) {
1483 // If we're accessing an SVE vector with split SVE objects...
1484 // - From the FP we need to move down past the PPR area:
1485 FPOffset -= PPRStackSize;
1486 // - From the SP we only need to move up to the ZPR area:
1487 SPOffset -= PPRStackSize;
1488 // Note: `SPOffset = SVEStackSize + ...`, so `-= PPRStackSize` results in
1489 // `SPOffset = ZPRStackSize + ...`.
1490 }
1491
1492 if (FPAfterSVECalleeSaves) {
1494 if (-ObjectOffset <= (int64_t)AFI->getSVECalleeSavedStackSize()) {
1497 }
1498 }
1499
1500 // Always use the FP for SVE spills if available and beneficial.
1501 if (hasFP(MF) && (SPOffset.getFixed() ||
1502 FPOffset.getScalable() < SPOffset.getScalable() ||
1503 RegInfo->hasStackRealignment(MF))) {
1504 FrameReg = RegInfo->getFrameRegister(MF);
1505 return FPOffset;
1506 }
1507 FrameReg = RegInfo->hasBasePointer(MF) ? RegInfo->getBaseRegister()
1508 : MCRegister(AArch64::SP);
1509
1510 return SPOffset;
1511 }
1512
1513 StackOffset SVEAreaOffset = {};
1514 if (FPAfterSVECalleeSaves) {
1515 // In this stack layout, the FP is in between the callee saves and other
1516 // SVE allocations.
1517 StackOffset SVECalleeSavedStack =
1519 if (UseFP) {
1520 if (isFixed)
1521 SVEAreaOffset = SVECalleeSavedStack;
1522 else if (!isCSR)
1523 SVEAreaOffset = SVECalleeSavedStack - SVEStackSize;
1524 } else {
1525 if (isFixed)
1526 SVEAreaOffset = SVEStackSize;
1527 else if (isCSR)
1528 SVEAreaOffset = SVEStackSize - SVECalleeSavedStack;
1529 }
1530 } else {
1531 if (UseFP && !(isFixed || isCSR))
1532 SVEAreaOffset = -SVEStackSize;
1533 if (!UseFP && (isFixed || isCSR))
1534 SVEAreaOffset = SVEStackSize;
1535 }
1536
1537 if (UseFP) {
1538 FrameReg = RegInfo->getFrameRegister(MF);
1539 return StackOffset::getFixed(FPOffset) + SVEAreaOffset;
1540 }
1541
1542 // Use the base pointer if we have one.
1543 if (RegInfo->hasBasePointer(MF))
1544 FrameReg = RegInfo->getBaseRegister();
1545 else {
1546 assert(!MFI.hasVarSizedObjects() &&
1547 "Can't use SP when we have var sized objects.");
1548 FrameReg = AArch64::SP;
1549 // If we're using the red zone for this function, the SP won't actually
1550 // be adjusted, so the offsets will be negative. They're also all
1551 // within range of the signed 9-bit immediate instructions.
1552 if (canUseRedZone(MF))
1553 Offset -= AFI->getLocalStackSize();
1554 }
1555
1556 return StackOffset::getFixed(Offset) + SVEAreaOffset;
1557}
1558
1560 // Do not set a kill flag on values that are also marked as live-in. This
1561 // happens with the @llvm-returnaddress intrinsic and with arguments passed in
1562 // callee saved registers.
1563 // Omitting the kill flags is conservatively correct even if the live-in
1564 // is not used after all.
1565 bool IsLiveIn = MF.getRegInfo().isLiveIn(Reg);
1566 return getKillRegState(!IsLiveIn);
1567}
1568
1570 MachineFunction &MF) {
1571 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1572 AttributeList Attrs = MF.getFunction().getAttributes();
1574 return Subtarget.isTargetMachO() &&
1575 !(Subtarget.getTargetLowering()->supportSwiftError() &&
1576 Attrs.hasAttrSomewhere(Attribute::SwiftError)) &&
1578 !AFL.requiresSaveVG(MF) && !AFI->isSVECC();
1579}
1580
1581static bool invalidateWindowsRegisterPairing(bool SpillExtendedVolatile,
1582 unsigned SpillCount, unsigned Reg1,
1583 unsigned Reg2, bool NeedsWinCFI,
1584 const TargetRegisterInfo *TRI) {
1585 // If we are generating register pairs for a Windows function that requires
1586 // EH support, then pair consecutive registers only. There are no unwind
1587 // opcodes for saves/restores of non-consecutive register pairs.
1588 // The unwind opcodes are save_regp, save_regp_x, save_fregp, save_frepg_x,
1589 // save_lrpair.
1590 // https://docs.microsoft.com/en-us/cpp/build/arm64-exception-handling
1591
1592 if (Reg2 == AArch64::FP)
1593 return true;
1594 if (!NeedsWinCFI)
1595 return false;
1596
1597 // ARM64EC introduced `save_any_regp`, which expects 16-byte alignment.
1598 // This is handled by only allowing paired spills for registers spilled at
1599 // even positions (which should be 16-byte aligned, as other GPRs/FPRs are
1600 // 8-bytes). We carve out an exception for {FP,LR}, which does not require
1601 // 16-byte alignment in the uop representation.
1602 if (TRI->getEncodingValue(Reg2) == TRI->getEncodingValue(Reg1) + 1)
1603 return SpillExtendedVolatile
1604 ? !((Reg1 == AArch64::FP && Reg2 == AArch64::LR) ||
1605 (SpillCount % 2) == 0)
1606 : false;
1607
1608 // If pairing a GPR with LR, the pair can be described by the save_lrpair
1609 // opcode. The save_lrpair opcode requires the first register to be odd.
1610 if (Reg1 >= AArch64::X19 && Reg1 <= AArch64::X27 &&
1611 (Reg1 - AArch64::X19) % 2 == 0 && Reg2 == AArch64::LR)
1612 return false;
1613 return true;
1614}
1615
1616/// Returns true if Reg1 and Reg2 cannot be paired using a ldp/stp instruction.
1617/// WindowsCFI requires that only consecutive registers can be paired.
1618/// LR and FP need to be allocated together when the frame needs to save
1619/// the frame-record. This means any other register pairing with LR is invalid.
1620static bool invalidateRegisterPairing(bool SpillExtendedVolatile,
1621 unsigned SpillCount, unsigned Reg1,
1622 unsigned Reg2, bool UsesWinAAPCS,
1623 bool NeedsWinCFI, bool NeedsFrameRecord,
1624 const TargetRegisterInfo *TRI) {
1625 if (UsesWinAAPCS)
1626 return invalidateWindowsRegisterPairing(SpillExtendedVolatile, SpillCount,
1627 Reg1, Reg2, NeedsWinCFI, TRI);
1628
1629 // If we need to store the frame record, don't pair any register
1630 // with LR other than FP.
1631 if (NeedsFrameRecord)
1632 return Reg2 == AArch64::LR;
1633
1634 return false;
1635}
1636
1637namespace {
1638
1639struct RegPairInfo {
1640 Register Reg1;
1641 Register Reg2;
1642 int FrameIdx;
1643 int Offset;
1644 enum RegType { GPR, FPR64, FPR128, PPR, ZPR, VG } Type;
1645 const TargetRegisterClass *RC;
1646
1647 RegPairInfo() = default;
1648
1649 bool isPaired() const { return Reg2.isValid(); }
1650
1651 bool isScalable() const { return Type == PPR || Type == ZPR; }
1652};
1653
1654} // end anonymous namespace
1655
1657 for (unsigned PReg = AArch64::P8; PReg <= AArch64::P15; ++PReg) {
1658 if (SavedRegs.test(PReg)) {
1659 unsigned PNReg = PReg - AArch64::P0 + AArch64::PN0;
1660 return MCRegister(PNReg);
1661 }
1662 }
1663 return MCRegister();
1664}
1665
1666// The multivector LD/ST are available only for SME or SVE2p1 targets
1668 MachineFunction &MF) {
1670 return false;
1671
1672 SMEAttrs FuncAttrs = MF.getInfo<AArch64FunctionInfo>()->getSMEFnAttrs();
1673 bool IsLocallyStreaming =
1674 FuncAttrs.hasStreamingBody() && !FuncAttrs.hasStreamingInterface();
1675
1676 // Only when in streaming mode SME2 instructions can be safely used.
1677 // It is not safe to use SME2 instructions when in streaming compatible or
1678 // locally streaming mode.
1679 return Subtarget.hasSVE2p1() ||
1680 (Subtarget.hasSME2() &&
1681 (!IsLocallyStreaming && Subtarget.isStreaming()));
1682}
1683
1685 MachineFunction &MF,
1687 const TargetRegisterInfo *TRI,
1689 bool NeedsFrameRecord) {
1690
1691 if (CSI.empty())
1692 return;
1693
1694 bool IsWindows = isTargetWindows(MF);
1696 unsigned StackHazardSize = getStackHazardSize(MF);
1697 MachineFrameInfo &MFI = MF.getFrameInfo();
1699 unsigned Count = CSI.size();
1700 (void)CC;
1701 // MachO's compact unwind format relies on all registers being stored in
1702 // pairs.
1703 assert((!produceCompactUnwindFrame(AFL, MF) ||
1706 (Count & 1) == 0) &&
1707 "Odd number of callee-saved regs to spill!");
1708 int ByteOffset = AFI->getCalleeSavedStackSize();
1709 int StackFillDir = -1;
1710 int RegInc = 1;
1711 unsigned FirstReg = 0;
1712 if (IsWindows) {
1713 // For WinCFI, fill the stack from the bottom up.
1714 ByteOffset = 0;
1715 StackFillDir = 1;
1716 // As the CSI array is reversed to match PrologEpilogInserter, iterate
1717 // backwards, to pair up registers starting from lower numbered registers.
1718 RegInc = -1;
1719 FirstReg = Count - 1;
1720 }
1721
1722 bool FPAfterSVECalleeSaves = AFL.hasSVECalleeSavesAboveFrameRecord(MF);
1723 // Windows AAPCS has x9-x15 as volatile registers, x16-x17 as intra-procedural
1724 // scratch, x18 as platform reserved. However, clang has extended calling
1725 // convensions such as preserve_most and preserve_all which treat these as
1726 // CSR. As such, the ARM64 unwind uOPs bias registers by 19. We use ARM64EC
1727 // uOPs which have separate restrictions. We need to check for that.
1728 //
1729 // NOTE: we currently do not account for the D registers as LLVM does not
1730 // support non-ABI compliant D register spills.
1731 bool SpillExtendedVolatile =
1732 IsWindows && llvm::any_of(CSI, [](const CalleeSavedInfo &CSI) {
1733 const auto &Reg = CSI.getReg();
1734 return Reg >= AArch64::X0 && Reg <= AArch64::X18;
1735 });
1736
1737 int ZPRByteOffset = 0;
1738 int PPRByteOffset = 0;
1739 bool SplitPPRs = AFI->hasSplitSVEObjects();
1740 if (SplitPPRs) {
1741 ZPRByteOffset = AFI->getZPRCalleeSavedStackSize();
1742 PPRByteOffset = AFI->getPPRCalleeSavedStackSize();
1743 } else if (!FPAfterSVECalleeSaves) {
1744 ZPRByteOffset =
1746 // Unused: Everything goes in ZPR space.
1747 PPRByteOffset = 0;
1748 }
1749
1750 bool NeedGapToAlignStack = AFI->hasCalleeSaveStackFreeSpace();
1751 Register LastReg = 0;
1752 bool HasCSHazardPadding = AFI->hasStackHazardSlotIndex() && !SplitPPRs;
1753
1754 auto AlignOffset = [StackFillDir](int Offset, int Align) {
1755 if (StackFillDir < 0)
1756 return alignDown(Offset, Align);
1757 return alignTo(Offset, Align);
1758 };
1759
1760 // When iterating backwards, the loop condition relies on unsigned wraparound.
1761 for (unsigned i = FirstReg; i < Count; i += RegInc) {
1762 RegPairInfo RPI;
1763 RPI.Reg1 = CSI[i].getReg();
1764
1765 if (AArch64::GPR64RegClass.contains(RPI.Reg1)) {
1766 RPI.Type = RegPairInfo::GPR;
1767 RPI.RC = &AArch64::GPR64RegClass;
1768 } else if (AArch64::FPR64RegClass.contains(RPI.Reg1)) {
1769 RPI.Type = RegPairInfo::FPR64;
1770 RPI.RC = &AArch64::FPR64RegClass;
1771 } else if (AArch64::FPR128RegClass.contains(RPI.Reg1)) {
1772 RPI.Type = RegPairInfo::FPR128;
1773 RPI.RC = &AArch64::FPR128RegClass;
1774 } else if (AArch64::ZPRRegClass.contains(RPI.Reg1)) {
1775 RPI.Type = RegPairInfo::ZPR;
1776 RPI.RC = &AArch64::ZPRRegClass;
1777 } else if (AArch64::PPRRegClass.contains(RPI.Reg1)) {
1778 RPI.Type = RegPairInfo::PPR;
1779 RPI.RC = &AArch64::PPRRegClass;
1780 } else if (RPI.Reg1 == AArch64::VG) {
1781 RPI.Type = RegPairInfo::VG;
1782 RPI.RC = &AArch64::FIXED_REGSRegClass;
1783 } else {
1784 llvm_unreachable("Unsupported register class.");
1785 }
1786
1787 int &ScalableByteOffset = RPI.Type == RegPairInfo::PPR && SplitPPRs
1788 ? PPRByteOffset
1789 : ZPRByteOffset;
1790
1791 // Add the stack hazard size as we transition from GPR->FPR CSRs.
1792 if (HasCSHazardPadding &&
1793 (!LastReg || !AArch64InstrInfo::isFpOrNEON(LastReg)) &&
1795 ByteOffset += StackFillDir * StackHazardSize;
1796 LastReg = RPI.Reg1;
1797
1798 bool NeedsWinCFI = AFL.needsWinCFI(MF);
1799 int Scale = TRI->getSpillSize(*RPI.RC);
1800 // Add the next reg to the pair if it is in the same register class.
1801 if (unsigned(i + RegInc) < Count && !HasCSHazardPadding) {
1802 MCRegister NextReg = CSI[i + RegInc].getReg();
1803 unsigned SpillCount = NeedsWinCFI ? FirstReg - i : i;
1804 int Aligned = AlignOffset(ByteOffset, Scale);
1805 int PairOffset = IsWindows ? Aligned : Aligned + StackFillDir * 2 * Scale;
1806 bool PairFitsImmRange =
1807 PairOffset / Scale >= -64 && PairOffset / Scale <= 63;
1808 switch (RPI.Type) {
1809 case RegPairInfo::GPR:
1810 if (AArch64::GPR64RegClass.contains(NextReg) && PairFitsImmRange &&
1811 !invalidateRegisterPairing(SpillExtendedVolatile, SpillCount,
1812 RPI.Reg1, NextReg, IsWindows,
1813 NeedsWinCFI, NeedsFrameRecord, TRI))
1814 RPI.Reg2 = NextReg;
1815 break;
1816 case RegPairInfo::FPR64:
1817 if (AArch64::FPR64RegClass.contains(NextReg) && PairFitsImmRange &&
1818 !invalidateRegisterPairing(SpillExtendedVolatile, SpillCount,
1819 RPI.Reg1, NextReg, IsWindows,
1820 NeedsWinCFI, NeedsFrameRecord, TRI))
1821 RPI.Reg2 = NextReg;
1822 break;
1823 case RegPairInfo::FPR128:
1824 if (AArch64::FPR128RegClass.contains(NextReg) && PairFitsImmRange)
1825 RPI.Reg2 = NextReg;
1826 break;
1827 case RegPairInfo::PPR:
1828 break;
1829 case RegPairInfo::ZPR:
1830 if (AFI->getPredicateRegForFillSpill() != 0 &&
1831 ((RPI.Reg1 - AArch64::Z0) & 1) == 0 && (NextReg == RPI.Reg1 + 1)) {
1832 // Calculate offset of register pair to see if pair instruction can be
1833 // used.
1834 int Offset = (ScalableByteOffset + StackFillDir * 2 * Scale) / Scale;
1835 if ((-16 <= Offset && Offset <= 14) && (Offset % 2 == 0))
1836 RPI.Reg2 = NextReg;
1837 }
1838 break;
1839 case RegPairInfo::VG:
1840 break;
1841 }
1842 }
1843
1844 // GPRs and FPRs are saved in pairs of 64-bit regs. We expect the CSI
1845 // list to come in sorted by frame index so that we can issue the store
1846 // pair instructions directly. Assert if we see anything otherwise.
1847 //
1848 // The order of the registers in the list is controlled by
1849 // getCalleeSavedRegs(), so they will always be in-order, as well.
1850 assert((!RPI.isPaired() ||
1851 (CSI[i].getFrameIdx() + RegInc == CSI[i + RegInc].getFrameIdx())) &&
1852 "Out of order callee saved regs!");
1853
1854 assert((!RPI.isPaired() || !NeedsFrameRecord || RPI.Reg2 != AArch64::FP ||
1855 RPI.Reg1 == AArch64::LR) &&
1856 "FrameRecord must be allocated together with LR");
1857
1858 // Windows AAPCS has FP and LR reversed.
1859 assert((!RPI.isPaired() || !NeedsFrameRecord || RPI.Reg1 != AArch64::FP ||
1860 RPI.Reg2 == AArch64::LR) &&
1861 "FrameRecord must be allocated together with LR");
1862
1863 // MachO's compact unwind format relies on all registers being stored in
1864 // adjacent register pairs.
1865 assert((!produceCompactUnwindFrame(AFL, MF) ||
1868 (RPI.isPaired() &&
1869 ((RPI.Reg1 == AArch64::LR && RPI.Reg2 == AArch64::FP) ||
1870 RPI.Reg1 + 1 == RPI.Reg2))) &&
1871 "Callee-save registers not saved as adjacent register pair!");
1872
1873 RPI.FrameIdx = CSI[i].getFrameIdx();
1874 if (IsWindows &&
1875 RPI.isPaired()) // RPI.FrameIdx must be the lower index of the pair
1876 RPI.FrameIdx = CSI[i + RegInc].getFrameIdx();
1877
1878 // Realign the scalable offset if necessary. This is relevant when spilling
1879 // predicates on Windows.
1880 if (RPI.isScalable() && ScalableByteOffset % Scale != 0)
1881 ScalableByteOffset = AlignOffset(ScalableByteOffset, Scale);
1882
1883 // Realign the fixed offset if necessary. This is relevant when spilling Q
1884 // registers after spilling an odd amount of X registers.
1885 if (!RPI.isScalable() && ByteOffset % Scale != 0)
1886 ByteOffset = AlignOffset(ByteOffset, Scale);
1887
1888 int OffsetPre = RPI.isScalable() ? ScalableByteOffset : ByteOffset;
1889 assert(OffsetPre % Scale == 0);
1890
1891 if (RPI.isScalable())
1892 ScalableByteOffset += StackFillDir * (RPI.isPaired() ? 2 * Scale : Scale);
1893 else
1894 ByteOffset += StackFillDir * (RPI.isPaired() ? 2 * Scale : Scale);
1895
1896 // Swift's async context is directly before FP, so allocate an extra
1897 // 8 bytes for it.
1898 if (NeedsFrameRecord && AFI->hasSwiftAsyncContext() &&
1899 ((!IsWindows && RPI.Reg2 == AArch64::FP) ||
1900 (IsWindows && RPI.Reg2 == AArch64::LR)))
1901 ByteOffset += StackFillDir * 8;
1902
1903 // Round up size of non-pair to pair size if we need to pad the
1904 // callee-save area to ensure 16-byte alignment.
1905 if (NeedGapToAlignStack && !IsWindows && !RPI.isScalable() &&
1906 RPI.Type != RegPairInfo::FPR128 && !RPI.isPaired() &&
1907 ByteOffset % 16 != 0) {
1908 ByteOffset += 8 * StackFillDir;
1909 assert(MFI.getObjectAlign(RPI.FrameIdx) <= Align(16));
1910 // A stack frame with a gap looks like this, bottom up:
1911 // d9, d8. x21, gap, x20, x19.
1912 // Set extra alignment on the x21 object to create the gap above it.
1913 MFI.setObjectAlignment(RPI.FrameIdx, Align(16));
1914 NeedGapToAlignStack = false;
1915 }
1916
1917 int OffsetPost = RPI.isScalable() ? ScalableByteOffset : ByteOffset;
1918 assert(OffsetPost % Scale == 0);
1919 // If filling top down (default), we want the offset after incrementing it.
1920 // If filling bottom up (WinCFI) we need the original offset.
1921 int Offset = IsWindows ? OffsetPre : OffsetPost;
1922
1923 // The FP, LR pair goes 8 bytes into our expanded 24-byte slot so that the
1924 // Swift context can directly precede FP.
1925 if (NeedsFrameRecord && AFI->hasSwiftAsyncContext() &&
1926 ((!IsWindows && RPI.Reg2 == AArch64::FP) ||
1927 (IsWindows && RPI.Reg2 == AArch64::LR)))
1928 Offset += 8;
1929 RPI.Offset = Offset / Scale;
1930
1931 assert((!RPI.isPaired() ||
1932 (!RPI.isScalable() && RPI.Offset >= -64 && RPI.Offset <= 63) ||
1933 (RPI.isScalable() && RPI.Offset >= -256 && RPI.Offset <= 255)) &&
1934 "Offset out of bounds for LDP/STP immediate");
1935
1936 auto isFrameRecord = [&] {
1937 if (RPI.isPaired())
1938 return IsWindows ? RPI.Reg1 == AArch64::FP && RPI.Reg2 == AArch64::LR
1939 : RPI.Reg1 == AArch64::LR && RPI.Reg2 == AArch64::FP;
1940 // Otherwise, look for the frame record as two unpaired registers. This is
1941 // needed for -aarch64-stack-hazard-size=<val>, which disables register
1942 // pairing (as the padding may be too large for the LDP/STP offset). Note:
1943 // On Windows, this check works out as current reg == FP, next reg == LR,
1944 // and on other platforms current reg == FP, previous reg == LR. This
1945 // works out as the correct pre-increment or post-increment offsets
1946 // respectively.
1947 return i > 0 && RPI.Reg1 == AArch64::FP &&
1948 CSI[i - 1].getReg() == AArch64::LR;
1949 };
1950
1951 // Save the offset to frame record so that the FP register can point to the
1952 // innermost frame record (spilled FP and LR registers).
1953 if (NeedsFrameRecord && isFrameRecord())
1955
1956 RegPairs.push_back(RPI);
1957 if (RPI.isPaired())
1958 i += RegInc;
1959 }
1960 if (IsWindows) {
1961 // If we need an alignment gap in the stack, align the topmost stack
1962 // object. A stack frame with a gap looks like this, bottom up:
1963 // x19, d8. d9, gap.
1964 // Set extra alignment on the topmost stack object (the first element in
1965 // CSI, which goes top down), to create the gap above it.
1966 if (AFI->hasCalleeSaveStackFreeSpace())
1967 MFI.setObjectAlignment(CSI[0].getFrameIdx(), Align(16));
1968 // We iterated bottom up over the registers; flip RegPairs back to top
1969 // down order.
1970 std::reverse(RegPairs.begin(), RegPairs.end());
1971 }
1972}
1973
1977 MachineFunction &MF = *MBB.getParent();
1978 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
1979 auto &TLI = *Subtarget.getTargetLowering();
1980 const AArch64InstrInfo &TII = *Subtarget.getInstrInfo();
1981 bool NeedsWinCFI = needsWinCFI(MF);
1982 DebugLoc DL;
1984
1985 computeCalleeSaveRegisterPairs(*this, MF, CSI, TRI, RegPairs, hasFP(MF));
1986
1987 MachineRegisterInfo &MRI = MF.getRegInfo();
1988 // Refresh the reserved regs in case there are any potential changes since the
1989 // last freeze.
1990 MRI.freezeReservedRegs();
1991
1992 if (homogeneousPrologEpilog(MF)) {
1993 auto MIB = BuildMI(MBB, MI, DL, TII.get(AArch64::HOM_Prolog))
1995
1996 for (auto &RPI : RegPairs) {
1997 MIB.addReg(RPI.Reg1);
1998 MIB.addReg(RPI.Reg2);
1999
2000 // Update register live in.
2001 if (!MRI.isReserved(RPI.Reg1))
2002 MBB.addLiveIn(RPI.Reg1);
2003 if (RPI.isPaired() && !MRI.isReserved(RPI.Reg2))
2004 MBB.addLiveIn(RPI.Reg2);
2005 }
2006 return true;
2007 }
2008 bool PTrueCreated = false;
2009 for (const RegPairInfo &RPI : llvm::reverse(RegPairs)) {
2010 Register Reg1 = RPI.Reg1;
2011 Register Reg2 = RPI.Reg2;
2012 unsigned StrOpc;
2013
2014 // Issue sequence of spills for cs regs. The first spill may be converted
2015 // to a pre-decrement store later by emitPrologue if the callee-save stack
2016 // area allocation can't be combined with the local stack area allocation.
2017 // For example:
2018 // stp x22, x21, [sp, #0] // addImm(+0)
2019 // stp x20, x19, [sp, #16] // addImm(+2)
2020 // stp fp, lr, [sp, #32] // addImm(+4)
2021 // Rationale: This sequence saves uop updates compared to a sequence of
2022 // pre-increment spills like stp xi,xj,[sp,#-16]!
2023 // Note: Similar rationale and sequence for restores in epilog.
2024 unsigned Size = TRI->getSpillSize(*RPI.RC);
2025 Align Alignment = TRI->getSpillAlign(*RPI.RC);
2026 switch (RPI.Type) {
2027 case RegPairInfo::GPR:
2028 StrOpc = RPI.isPaired() ? AArch64::STPXi : AArch64::STRXui;
2029 break;
2030 case RegPairInfo::FPR64:
2031 StrOpc = RPI.isPaired() ? AArch64::STPDi : AArch64::STRDui;
2032 break;
2033 case RegPairInfo::FPR128:
2034 StrOpc = RPI.isPaired() ? AArch64::STPQi : AArch64::STRQui;
2035 break;
2036 case RegPairInfo::ZPR:
2037 StrOpc = RPI.isPaired() ? AArch64::ST1B_2Z_IMM : AArch64::STR_ZXI;
2038 break;
2039 case RegPairInfo::PPR:
2040 StrOpc = AArch64::STR_PXI;
2041 break;
2042 case RegPairInfo::VG:
2043 StrOpc = AArch64::STRXui;
2044 break;
2045 }
2046
2047 Register X0Scratch;
2048 llvm::scope_exit RestoreX0([&] {
2049 if (X0Scratch != AArch64::NoRegister)
2050 BuildMI(MBB, MI, DL, TII.get(TargetOpcode::COPY), AArch64::X0)
2051 .addReg(X0Scratch)
2053 });
2054
2055 if (Reg1 == AArch64::VG) {
2056 // Find an available register to store value of VG to.
2057 Reg1 = findScratchNonCalleeSaveRegister(&MBB, true);
2058 assert(Reg1 != AArch64::NoRegister);
2059 if (MF.getSubtarget<AArch64Subtarget>().hasSVE()) {
2060 BuildMI(MBB, MI, DL, TII.get(AArch64::CNTD_XPiI), Reg1)
2061 .addImm(31)
2062 .addImm(1)
2064 } else {
2066 if (any_of(MBB.liveins(),
2067 [&STI](const MachineBasicBlock::RegisterMaskPair &LiveIn) {
2068 return STI.getRegisterInfo()->isSuperOrSubRegisterEq(
2069 AArch64::X0, LiveIn.PhysReg);
2070 })) {
2071 X0Scratch = Reg1;
2072 BuildMI(MBB, MI, DL, TII.get(TargetOpcode::COPY), X0Scratch)
2073 .addReg(AArch64::X0)
2075 }
2076
2077 RTLIB::Libcall LC = RTLIB::SMEABI_GET_CURRENT_VG;
2078 const uint32_t *RegMask =
2079 TRI->getCallPreservedMask(MF, TLI.getLibcallCallingConv(LC));
2080 BuildMI(MBB, MI, DL, TII.get(AArch64::BL))
2081 .addExternalSymbol(TLI.getLibcallName(LC))
2082 .addRegMask(RegMask)
2083 .addReg(AArch64::X0, RegState::ImplicitDefine)
2085 Reg1 = AArch64::X0;
2086 }
2087 }
2088
2089 LLVM_DEBUG({
2090 dbgs() << "CSR spill: (" << printReg(Reg1, TRI);
2091 if (RPI.isPaired())
2092 dbgs() << ", " << printReg(Reg2, TRI);
2093 dbgs() << ") -> fi#(" << RPI.FrameIdx;
2094 if (RPI.isPaired())
2095 dbgs() << ", " << RPI.FrameIdx + 1;
2096 dbgs() << ")\n";
2097 });
2098
2099 assert((!isTargetWindows(MF) ||
2100 !(Reg1 == AArch64::LR && Reg2 == AArch64::FP)) &&
2101 "Windows unwdinding requires a consecutive (FP,LR) pair");
2102 // Windows unwind codes require consecutive registers if registers are
2103 // paired. Make the switch here, so that the code below will save (x,x+1)
2104 // and not (x+1,x).
2105 unsigned FrameIdxReg1 = RPI.FrameIdx;
2106 unsigned FrameIdxReg2 = RPI.FrameIdx + 1;
2107 if (isTargetWindows(MF) && RPI.isPaired()) {
2108 std::swap(Reg1, Reg2);
2109 std::swap(FrameIdxReg1, FrameIdxReg2);
2110 }
2111
2112 if (RPI.isPaired() && RPI.isScalable()) {
2113 [[maybe_unused]] const AArch64Subtarget &Subtarget =
2116 unsigned PnReg = AFI->getPredicateRegForFillSpill();
2117 assert((PnReg != 0 && enableMultiVectorSpillFill(Subtarget, MF)) &&
2118 "Expects SVE2.1 or SME2 target and a predicate register");
2119#ifdef EXPENSIVE_CHECKS
2120 auto IsPPR = [](const RegPairInfo &c) {
2121 return c.Reg1 == RegPairInfo::PPR;
2122 };
2123 auto PPRBegin = std::find_if(RegPairs.begin(), RegPairs.end(), IsPPR);
2124 auto IsZPR = [](const RegPairInfo &c) {
2125 return c.Type == RegPairInfo::ZPR;
2126 };
2127 auto ZPRBegin = std::find_if(RegPairs.begin(), RegPairs.end(), IsZPR);
2128 assert(!(PPRBegin < ZPRBegin) &&
2129 "Expected callee save predicate to be handled first");
2130#endif
2131 if (!PTrueCreated) {
2132 PTrueCreated = true;
2133 BuildMI(MBB, MI, DL, TII.get(AArch64::PTRUE_C_B), PnReg)
2135 }
2136 MachineInstrBuilder MIB = BuildMI(MBB, MI, DL, TII.get(StrOpc));
2137 if (!MRI.isReserved(Reg1))
2138 MBB.addLiveIn(Reg1);
2139 if (!MRI.isReserved(Reg2))
2140 MBB.addLiveIn(Reg2);
2141 MIB.addReg(/*PairRegs*/ AArch64::Z0_Z1 + (RPI.Reg1 - AArch64::Z0));
2143 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2144 MachineMemOperand::MOStore, Size, Alignment));
2145 MIB.addReg(PnReg);
2146 MIB.addReg(AArch64::SP)
2147 .addImm(RPI.Offset / 2) // [sp, #imm*2*vscale],
2148 // where 2*vscale is implicit
2151 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2152 MachineMemOperand::MOStore, Size, Alignment));
2153 if (NeedsWinCFI)
2154 insertSEH(MIB, TII, MachineInstr::FrameSetup);
2155 } else { // The code when the pair of ZReg is not present
2156 MachineInstrBuilder MIB = BuildMI(MBB, MI, DL, TII.get(StrOpc));
2157 if (!MRI.isReserved(Reg1))
2158 MBB.addLiveIn(Reg1);
2159 if (RPI.isPaired()) {
2160 if (!MRI.isReserved(Reg2))
2161 MBB.addLiveIn(Reg2);
2162 MIB.addReg(Reg2, getPrologueDeath(MF, Reg2));
2164 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2165 MachineMemOperand::MOStore, Size, Alignment));
2166 }
2167 MIB.addReg(Reg1, getPrologueDeath(MF, Reg1))
2168 .addReg(AArch64::SP)
2169 .addImm(RPI.Offset) // [sp, #offset*vscale],
2170 // where factor*vscale is implicit
2173 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2174 MachineMemOperand::MOStore, Size, Alignment));
2175 if (NeedsWinCFI)
2176 insertSEH(MIB, TII, MachineInstr::FrameSetup);
2177 }
2178 // Update the StackIDs of the SVE stack slots.
2179 MachineFrameInfo &MFI = MF.getFrameInfo();
2180 if (RPI.Type == RegPairInfo::ZPR) {
2181 MFI.setStackID(FrameIdxReg1, TargetStackID::ScalableVector);
2182 if (RPI.isPaired())
2183 MFI.setStackID(FrameIdxReg2, TargetStackID::ScalableVector);
2184 } else if (RPI.Type == RegPairInfo::PPR) {
2186 if (RPI.isPaired())
2188 }
2189 }
2190 return true;
2191}
2192
2196 MachineFunction &MF = *MBB.getParent();
2197 const AArch64InstrInfo &TII =
2198 *MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
2199 DebugLoc DL;
2201 bool NeedsWinCFI = needsWinCFI(MF);
2202
2203 if (MBBI != MBB.end())
2204 DL = MBBI->getDebugLoc();
2205
2206 computeCalleeSaveRegisterPairs(*this, MF, CSI, TRI, RegPairs, hasFP(MF));
2207 if (homogeneousPrologEpilog(MF, &MBB)) {
2208 auto MIB = BuildMI(MBB, MBBI, DL, TII.get(AArch64::HOM_Epilog))
2210 for (auto &RPI : RegPairs) {
2211 MIB.addReg(RPI.Reg1, RegState::Define);
2212 MIB.addReg(RPI.Reg2, RegState::Define);
2213 }
2214 return true;
2215 }
2216
2217 // For performance reasons restore SVE register in increasing order
2218 auto IsPPR = [](const RegPairInfo &c) { return c.Type == RegPairInfo::PPR; };
2219 auto PPRBegin = llvm::find_if(RegPairs, IsPPR);
2220 auto PPREnd = std::find_if_not(PPRBegin, RegPairs.end(), IsPPR);
2221 std::reverse(PPRBegin, PPREnd);
2222 auto IsZPR = [](const RegPairInfo &c) { return c.Type == RegPairInfo::ZPR; };
2223 auto ZPRBegin = llvm::find_if(RegPairs, IsZPR);
2224 auto ZPREnd = std::find_if_not(ZPRBegin, RegPairs.end(), IsZPR);
2225 std::reverse(ZPRBegin, ZPREnd);
2226
2227 bool PTrueCreated = false;
2228 for (const RegPairInfo &RPI : RegPairs) {
2229 Register Reg1 = RPI.Reg1;
2230 Register Reg2 = RPI.Reg2;
2231
2232 // Issue sequence of restores for cs regs. The last restore may be converted
2233 // to a post-increment load later by emitEpilogue if the callee-save stack
2234 // area allocation can't be combined with the local stack area allocation.
2235 // For example:
2236 // ldp fp, lr, [sp, #32] // addImm(+4)
2237 // ldp x20, x19, [sp, #16] // addImm(+2)
2238 // ldp x22, x21, [sp, #0] // addImm(+0)
2239 // Note: see comment in spillCalleeSavedRegisters()
2240 unsigned LdrOpc;
2241 unsigned Size = TRI->getSpillSize(*RPI.RC);
2242 Align Alignment = TRI->getSpillAlign(*RPI.RC);
2243 switch (RPI.Type) {
2244 case RegPairInfo::GPR:
2245 LdrOpc = RPI.isPaired() ? AArch64::LDPXi : AArch64::LDRXui;
2246 break;
2247 case RegPairInfo::FPR64:
2248 LdrOpc = RPI.isPaired() ? AArch64::LDPDi : AArch64::LDRDui;
2249 break;
2250 case RegPairInfo::FPR128:
2251 LdrOpc = RPI.isPaired() ? AArch64::LDPQi : AArch64::LDRQui;
2252 break;
2253 case RegPairInfo::ZPR:
2254 LdrOpc = RPI.isPaired() ? AArch64::LD1B_2Z_IMM : AArch64::LDR_ZXI;
2255 break;
2256 case RegPairInfo::PPR:
2257 LdrOpc = AArch64::LDR_PXI;
2258 break;
2259 case RegPairInfo::VG:
2260 continue;
2261 }
2262 LLVM_DEBUG({
2263 dbgs() << "CSR restore: (" << printReg(Reg1, TRI);
2264 if (RPI.isPaired())
2265 dbgs() << ", " << printReg(Reg2, TRI);
2266 dbgs() << ") -> fi#(" << RPI.FrameIdx;
2267 if (RPI.isPaired())
2268 dbgs() << ", " << RPI.FrameIdx + 1;
2269 dbgs() << ")\n";
2270 });
2271
2272 // Windows unwind codes require consecutive registers if registers are
2273 // paired. Make the switch here, so that the code below will save (x,x+1)
2274 // and not (x+1,x).
2275 unsigned FrameIdxReg1 = RPI.FrameIdx;
2276 unsigned FrameIdxReg2 = RPI.FrameIdx + 1;
2277 if (isTargetWindows(MF) && RPI.isPaired()) {
2278 std::swap(Reg1, Reg2);
2279 std::swap(FrameIdxReg1, FrameIdxReg2);
2280 }
2281
2283 if (RPI.isPaired() && RPI.isScalable()) {
2284 [[maybe_unused]] const AArch64Subtarget &Subtarget =
2286 unsigned PnReg = AFI->getPredicateRegForFillSpill();
2287 assert((PnReg != 0 && enableMultiVectorSpillFill(Subtarget, MF)) &&
2288 "Expects SVE2.1 or SME2 target and a predicate register");
2289#ifdef EXPENSIVE_CHECKS
2290 assert(!(PPRBegin < ZPRBegin) &&
2291 "Expected callee save predicate to be handled first");
2292#endif
2293 if (!PTrueCreated) {
2294 PTrueCreated = true;
2295 BuildMI(MBB, MBBI, DL, TII.get(AArch64::PTRUE_C_B), PnReg)
2297 }
2298 MachineInstrBuilder MIB = BuildMI(MBB, MBBI, DL, TII.get(LdrOpc));
2299 MIB.addReg(/*PairRegs*/ AArch64::Z0_Z1 + (RPI.Reg1 - AArch64::Z0),
2300 getDefRegState(true));
2302 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2303 MachineMemOperand::MOLoad, Size, Alignment));
2304 MIB.addReg(PnReg);
2305 MIB.addReg(AArch64::SP)
2306 .addImm(RPI.Offset / 2) // [sp, #imm*2*vscale]
2307 // where 2*vscale is implicit
2310 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2311 MachineMemOperand::MOLoad, Size, Alignment));
2312 if (NeedsWinCFI)
2313 insertSEH(MIB, TII, MachineInstr::FrameDestroy);
2314 } else {
2315 MachineInstrBuilder MIB = BuildMI(MBB, MBBI, DL, TII.get(LdrOpc));
2316 if (RPI.isPaired()) {
2317 MIB.addReg(Reg2, getDefRegState(true));
2319 MachinePointerInfo::getFixedStack(MF, FrameIdxReg2),
2320 MachineMemOperand::MOLoad, Size, Alignment));
2321 }
2322 MIB.addReg(Reg1, getDefRegState(true));
2323 MIB.addReg(AArch64::SP)
2324 .addImm(RPI.Offset) // [sp, #offset*vscale]
2325 // where factor*vscale is implicit
2328 MachinePointerInfo::getFixedStack(MF, FrameIdxReg1),
2329 MachineMemOperand::MOLoad, Size, Alignment));
2330 if (NeedsWinCFI)
2331 insertSEH(MIB, TII, MachineInstr::FrameDestroy);
2332 }
2333 }
2334 return true;
2335}
2336
2337// Return the FrameID for a MMO.
2338static std::optional<int> getMMOFrameID(MachineMemOperand *MMO,
2339 const MachineFrameInfo &MFI) {
2340 auto *PSV =
2342 if (PSV)
2343 return std::optional<int>(PSV->getFrameIndex());
2344
2345 if (MMO->getValue()) {
2346 if (auto *Al = dyn_cast<AllocaInst>(getUnderlyingObject(MMO->getValue()))) {
2347 for (int FI = MFI.getObjectIndexBegin(); FI < MFI.getObjectIndexEnd();
2348 FI++)
2349 if (MFI.getObjectAllocation(FI) == Al)
2350 return FI;
2351 }
2352 }
2353
2354 return std::nullopt;
2355}
2356
2357// Return the FrameID for a Load/Store instruction by looking at the first MMO.
2358static std::optional<int> getLdStFrameID(const MachineInstr &MI,
2359 const MachineFrameInfo &MFI) {
2360 if (!MI.mayLoadOrStore() || MI.getNumMemOperands() < 1)
2361 return std::nullopt;
2362
2363 return getMMOFrameID(*MI.memoperands_begin(), MFI);
2364}
2365
2366// Returns true if the LDST MachineInstr \p MI is a PPR access.
2367static bool isPPRAccess(const MachineInstr &MI) {
2368 return AArch64::PPRRegClass.contains(MI.getOperand(0).getReg());
2369}
2370
2371// Check if a Hazard slot is needed for the current function, and if so create
2372// one for it. The index is stored in AArch64FunctionInfo->StackHazardSlotIndex,
2373// which can be used to determine if any hazard padding is needed.
2374void AArch64FrameLowering::determineStackHazardSlot(
2375 MachineFunction &MF, BitVector &SavedRegs) const {
2376 unsigned StackHazardSize = getStackHazardSize(MF);
2377 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
2378 if (StackHazardSize == 0 || StackHazardSize % 16 != 0 ||
2380 return;
2381
2382 // Stack hazards are only needed in streaming functions.
2383 SMEAttrs Attrs = AFI->getSMEFnAttrs();
2384 if (!StackHazardInNonStreaming && Attrs.hasNonStreamingInterfaceAndBody())
2385 return;
2386
2387 MachineFrameInfo &MFI = MF.getFrameInfo();
2388
2389 // Add a hazard slot if there are any CSR FPR registers, or are any fp-only
2390 // stack objects.
2391 bool HasFPRCSRs = any_of(SavedRegs.set_bits(), [](unsigned Reg) {
2392 return AArch64::FPR64RegClass.contains(Reg) ||
2393 AArch64::FPR128RegClass.contains(Reg) ||
2394 AArch64::ZPRRegClass.contains(Reg);
2395 });
2396 bool HasPPRCSRs = any_of(SavedRegs.set_bits(), [](unsigned Reg) {
2397 return AArch64::PPRRegClass.contains(Reg);
2398 });
2399 bool HasFPRStackObjects = false;
2400 bool HasPPRStackObjects = false;
2401 if (!HasFPRCSRs || SplitSVEObjects) {
2402 enum SlotType : uint8_t {
2403 Unknown = 0,
2404 ZPRorFPR = 1 << 0,
2405 PPR = 1 << 1,
2406 GPR = 1 << 2,
2408 };
2409
2410 // Find stack slots solely used for one kind of register (ZPR, PPR, etc.),
2411 // based on the kinds of accesses used in the function.
2412 SmallVector<SlotType> SlotTypes(MFI.getObjectIndexEnd(), SlotType::Unknown);
2413 for (auto &MBB : MF) {
2414 for (auto &MI : MBB) {
2415 std::optional<int> FI = getLdStFrameID(MI, MFI);
2416 if (!FI || FI < 0 || FI > int(SlotTypes.size()))
2417 continue;
2418 if (MFI.hasScalableStackID(*FI)) {
2419 SlotTypes[*FI] |=
2420 isPPRAccess(MI) ? SlotType::PPR : SlotType::ZPRorFPR;
2421 } else {
2422 SlotTypes[*FI] |= AArch64InstrInfo::isFpOrNEON(MI)
2423 ? SlotType::ZPRorFPR
2424 : SlotType::GPR;
2425 }
2426 }
2427 }
2428
2429 for (int FI = 0; FI < int(SlotTypes.size()); ++FI) {
2430 HasFPRStackObjects |= SlotTypes[FI] == SlotType::ZPRorFPR;
2431 // For SplitSVEObjects remember that this stack slot is a predicate, this
2432 // will be needed later when determining the frame layout.
2433 if (SlotTypes[FI] == SlotType::PPR) {
2435 HasPPRStackObjects = true;
2436 }
2437 }
2438 }
2439
2440 if (HasFPRCSRs || HasFPRStackObjects) {
2441 int ID = MFI.CreateStackObject(StackHazardSize, Align(16), false);
2442 LLVM_DEBUG(dbgs() << "Created Hazard slot at " << ID << " size "
2443 << StackHazardSize << "\n");
2444 AFI->setStackHazardSlotIndex(ID);
2445 }
2446
2447 if (!AFI->hasStackHazardSlotIndex())
2448 return;
2449
2450 if (SplitSVEObjects) {
2451 CallingConv::ID CC = MF.getFunction().getCallingConv();
2452 if (AFI->isSVECC() || CC == CallingConv::AArch64_SVE_VectorCall) {
2453 AFI->setSplitSVEObjects(true);
2454 LLVM_DEBUG(dbgs() << "Using SplitSVEObjects for SVE CC function\n");
2455 return;
2456 }
2457
2458 // We only use SplitSVEObjects in non-SVE CC functions if there's a
2459 // possibility of a stack hazard between PPRs and ZPRs/FPRs.
2460 LLVM_DEBUG(dbgs() << "Determining if SplitSVEObjects should be used in "
2461 "non-SVE CC function...\n");
2462
2463 // If another calling convention is explicitly set FPRs can't be promoted to
2464 // ZPR callee-saves.
2466 LLVM_DEBUG(
2467 dbgs()
2468 << "Calling convention is not supported with SplitSVEObjects\n");
2469 return;
2470 }
2471
2472 if (!HasPPRCSRs && !HasPPRStackObjects) {
2473 LLVM_DEBUG(
2474 dbgs() << "Not using SplitSVEObjects as no PPRs are on the stack\n");
2475 return;
2476 }
2477
2478 if (!HasFPRCSRs && !HasFPRStackObjects) {
2479 LLVM_DEBUG(
2480 dbgs()
2481 << "Not using SplitSVEObjects as no FPRs or ZPRs are on the stack\n");
2482 return;
2483 }
2484
2485 [[maybe_unused]] const AArch64Subtarget &Subtarget =
2486 MF.getSubtarget<AArch64Subtarget>();
2488 "Expected SVE to be available for PPRs");
2489
2490 const TargetRegisterInfo *TRI = MF.getSubtarget().getRegisterInfo();
2491 // With SplitSVEObjects the CS hazard padding is placed between the
2492 // PPRs and ZPRs. If there are any FPR CS there would be a hazard between
2493 // them and the CS GRPs. Avoid this by promoting all FPR CS to ZPRs.
2494 BitVector FPRZRegs(SavedRegs.size());
2495 for (size_t Reg = 0, E = SavedRegs.size(); HasFPRCSRs && Reg < E; ++Reg) {
2496 BitVector::reference RegBit = SavedRegs[Reg];
2497 if (!RegBit)
2498 continue;
2499 unsigned SubRegIdx = 0;
2500 if (AArch64::FPR64RegClass.contains(Reg))
2501 SubRegIdx = AArch64::dsub;
2502 else if (AArch64::FPR128RegClass.contains(Reg))
2503 SubRegIdx = AArch64::zsub;
2504 else
2505 continue;
2506 // Clear the bit for the FPR save.
2507 RegBit = false;
2508 // Mark that we should save the corresponding ZPR.
2509 Register ZReg =
2510 TRI->getMatchingSuperReg(Reg, SubRegIdx, &AArch64::ZPRRegClass);
2511 FPRZRegs.set(ZReg);
2512 }
2513 SavedRegs |= FPRZRegs;
2514
2515 AFI->setSplitSVEObjects(true);
2516 LLVM_DEBUG(dbgs() << "SplitSVEObjects enabled!\n");
2517 }
2518}
2519
2521 BitVector &SavedRegs,
2522 RegScavenger *RS) const {
2523 // All calls are tail calls in GHC calling conv, and functions have no
2524 // prologue/epilogue.
2526 return;
2527
2528 const AArch64Subtarget &Subtarget = MF.getSubtarget<AArch64Subtarget>();
2529
2531 const AArch64RegisterInfo *RegInfo = Subtarget.getRegisterInfo();
2533 unsigned UnspilledCSGPR = AArch64::NoRegister;
2534 unsigned UnspilledCSGPRPaired = AArch64::NoRegister;
2535
2536 MachineFrameInfo &MFI = MF.getFrameInfo();
2537 const MCPhysReg *CSRegs = MF.getRegInfo().getCalleeSavedRegs();
2538
2539 MCRegister BasePointerReg =
2540 RegInfo->hasBasePointer(MF) ? RegInfo->getBaseRegister() : MCRegister();
2541
2542 unsigned ExtraCSSpill = 0;
2543 bool HasUnpairedGPR64 = false;
2544 bool HasPairZReg = false;
2545 BitVector UserReservedRegs = RegInfo->getUserReservedRegs(MF);
2546 BitVector ReservedRegs = RegInfo->getReservedRegs(MF);
2547
2548 // Figure out which callee-saved registers to save/restore.
2549 for (unsigned i = 0; CSRegs[i]; ++i) {
2550 const MCRegister Reg = CSRegs[i];
2551
2552 // Add the base pointer register to SavedRegs if it is callee-save.
2553 if (Reg == BasePointerReg)
2554 SavedRegs.set(Reg);
2555
2556 // Don't save manually reserved registers set through +reserve-x#i,
2557 // even for callee-saved registers, as per GCC's behavior.
2558 if (UserReservedRegs[Reg]) {
2559 SavedRegs.reset(Reg);
2560 continue;
2561 }
2562
2563 bool RegUsed = SavedRegs.test(Reg);
2564 MCRegister PairedReg;
2565 const bool RegIsGPR64 = AArch64::GPR64RegClass.contains(Reg);
2566 if (RegIsGPR64 || AArch64::FPR64RegClass.contains(Reg) ||
2567 AArch64::FPR128RegClass.contains(Reg)) {
2568 // Compensate for odd numbers of GP CSRs.
2569 // For now, all the known cases of odd number of CSRs are of GPRs.
2570 if (HasUnpairedGPR64)
2571 PairedReg = CSRegs[i % 2 == 0 ? i - 1 : i + 1];
2572 else
2573 PairedReg = CSRegs[i ^ 1];
2574 }
2575
2576 // If the function requires all the GP registers to save (SavedRegs),
2577 // and there are an odd number of GP CSRs at the same time (CSRegs),
2578 // PairedReg could be in a different register class from Reg, which would
2579 // lead to a FPR (usually D8) accidentally being marked saved.
2580 if (RegIsGPR64 && !AArch64::GPR64RegClass.contains(PairedReg)) {
2581 PairedReg = AArch64::NoRegister;
2582 HasUnpairedGPR64 = true;
2583 }
2584 assert(PairedReg == AArch64::NoRegister ||
2585 AArch64::GPR64RegClass.contains(Reg, PairedReg) ||
2586 AArch64::FPR64RegClass.contains(Reg, PairedReg) ||
2587 AArch64::FPR128RegClass.contains(Reg, PairedReg));
2588
2589 if (!RegUsed) {
2590 if (AArch64::GPR64RegClass.contains(Reg) && !ReservedRegs[Reg]) {
2591 UnspilledCSGPR = Reg;
2592 UnspilledCSGPRPaired = PairedReg;
2593 }
2594 continue;
2595 }
2596
2597 // MachO's compact unwind format relies on all registers being stored in
2598 // pairs.
2599 // FIXME: the usual format is actually better if unwinding isn't needed.
2600 if (producePairRegisters(MF) && PairedReg != AArch64::NoRegister &&
2601 !SavedRegs.test(PairedReg)) {
2602 SavedRegs.set(PairedReg);
2603 if (AArch64::GPR64RegClass.contains(PairedReg) &&
2604 !ReservedRegs[PairedReg])
2605 ExtraCSSpill = PairedReg;
2606 }
2607 // Check if there is a pair of ZRegs, so it can select PReg for spill/fill
2608 HasPairZReg |= (AArch64::ZPRRegClass.contains(Reg, CSRegs[i ^ 1]) &&
2609 SavedRegs.test(CSRegs[i ^ 1]));
2610 }
2611
2612 if (HasPairZReg && enableMultiVectorSpillFill(Subtarget, MF)) {
2614 // Find a suitable predicate register for the multi-vector spill/fill
2615 // instructions.
2616 MCRegister PnReg = findFreePredicateReg(SavedRegs);
2617 if (PnReg.isValid())
2618 AFI->setPredicateRegForFillSpill(PnReg);
2619 // If no free callee-save has been found assign one.
2620 if (!AFI->getPredicateRegForFillSpill() &&
2621 MF.getFunction().getCallingConv() ==
2623 SavedRegs.set(AArch64::P8);
2624 AFI->setPredicateRegForFillSpill(AArch64::PN8);
2625 }
2626
2627 assert(!ReservedRegs[AFI->getPredicateRegForFillSpill()] &&
2628 "Predicate cannot be a reserved register");
2629 }
2630
2632 !Subtarget.isTargetWindows()) {
2633 // For Windows calling convention on a non-windows OS, where X18 is treated
2634 // as reserved, back up X18 when entering non-windows code (marked with the
2635 // Windows calling convention) and restore when returning regardless of
2636 // whether the individual function uses it - it might call other functions
2637 // that clobber it.
2638 SavedRegs.set(AArch64::X18);
2639 }
2640
2641 // Determine if a Hazard slot should be used and where it should go.
2642 // If SplitSVEObjects is used, the hazard padding is placed between the PPRs
2643 // and ZPRs. Otherwise, it goes in the callee save area.
2644 determineStackHazardSlot(MF, SavedRegs);
2645
2646 // Calculates the callee saved stack size.
2647 unsigned CSStackSize = 0;
2648 unsigned ZPRCSStackSize = 0;
2649 unsigned PPRCSStackSize = 0;
2651 for (unsigned Reg : SavedRegs.set_bits()) {
2652 auto *RC = TRI->getMinimalPhysRegClass(MCRegister(Reg));
2653 assert(RC && "expected register class!");
2654 auto SpillSize = TRI->getSpillSize(*RC);
2655 bool IsZPR = AArch64::ZPRRegClass.contains(Reg);
2656 bool IsPPR = !IsZPR && AArch64::PPRRegClass.contains(Reg);
2657 if (IsZPR)
2658 ZPRCSStackSize += SpillSize;
2659 else if (IsPPR)
2660 PPRCSStackSize += SpillSize;
2661 else {
2662 // A register and its super-register can both appear in SavedRegs.
2663 // Only the widest register is actually spilled, so skip such
2664 // sub-registers here to avoid double-counting the overlap.
2665 bool SavedSuper = any_of(TRI->superregs(Reg), [&](MCPhysReg SuperReg) {
2666 return SavedRegs.test(SuperReg);
2667 });
2668 if (!SavedSuper)
2669 CSStackSize += SpillSize;
2670 }
2671 }
2672
2673 // Save number of saved regs, so we can easily update CSStackSize later to
2674 // account for any additional 64-bit GPR saves. Note: After this point
2675 // only 64-bit GPRs can be added to SavedRegs.
2676 unsigned NumSavedRegs = SavedRegs.count();
2677
2678 // If we have hazard padding in the CS area add that to the size.
2680 CSStackSize += getStackHazardSize(MF);
2681
2682 // Increase the callee-saved stack size if the function has streaming mode
2683 // changes, as we will need to spill the value of the VG register.
2684 if (requiresSaveVG(MF))
2685 CSStackSize += 8;
2686
2687 // If we must call __arm_get_current_vg in the prologue preserve the LR.
2688 if (requiresSaveVG(MF) && !Subtarget.hasSVE())
2689 SavedRegs.set(AArch64::LR);
2690
2691 // The frame record needs to be created by saving the appropriate registers
2692 uint64_t EstimatedStackSize = MFI.estimateStackSize(MF);
2693 if (hasFP(MF) ||
2694 windowsRequiresStackProbe(MF, EstimatedStackSize + CSStackSize + 16)) {
2695 SavedRegs.set(AArch64::FP);
2696 SavedRegs.set(AArch64::LR);
2697 }
2698
2699 LLVM_DEBUG({
2700 dbgs() << "*** determineCalleeSaves\nSaved CSRs:";
2701 for (unsigned Reg : SavedRegs.set_bits())
2702 dbgs() << ' ' << printReg(MCRegister(Reg), RegInfo);
2703 dbgs() << "\n";
2704 });
2705
2706 // If any callee-saved registers are used, the frame cannot be eliminated.
2707 auto [ZPRLocalStackSize, PPRLocalStackSize] =
2709 uint64_t SVELocals = ZPRLocalStackSize + PPRLocalStackSize;
2710 uint64_t SVEStackSize =
2711 alignTo(ZPRCSStackSize + PPRCSStackSize + SVELocals, 16);
2712 bool CanEliminateFrame = (SavedRegs.count() == 0) && !SVEStackSize;
2713
2714 // The CSR spill slots have not been allocated yet, so estimateStackSize
2715 // won't include them.
2716 unsigned EstimatedStackSizeLimit = estimateRSStackSizeLimit(MF);
2717
2718 // We may address some of the stack above the canonical frame address, either
2719 // for our own arguments or during a call. Include that in calculating whether
2720 // we have complicated addressing concerns.
2721 int64_t CalleeStackUsed = 0;
2722 for (int I = MFI.getObjectIndexBegin(); I != 0; ++I) {
2723 int64_t FixedOff = MFI.getObjectOffset(I);
2724 if (FixedOff > CalleeStackUsed)
2725 CalleeStackUsed = FixedOff;
2726 }
2727
2728 // Conservatively always assume BigStack when there are SVE spills.
2729 bool BigStack = SVEStackSize || (EstimatedStackSize + CSStackSize +
2730 CalleeStackUsed) > EstimatedStackSizeLimit;
2731 if (BigStack || !CanEliminateFrame || RegInfo->cannotEliminateFrame(MF))
2732 AFI->setHasStackFrame(true);
2733
2734 // Estimate if we might need to scavenge a register at some point in order
2735 // to materialize a stack offset. If so, either spill one additional
2736 // callee-saved register or reserve a special spill slot to facilitate
2737 // register scavenging. If we already spilled an extra callee-saved register
2738 // above to keep the number of spills even, we don't need to do anything else
2739 // here.
2740 if (BigStack) {
2741 if (!ExtraCSSpill && UnspilledCSGPR != AArch64::NoRegister) {
2742 LLVM_DEBUG(dbgs() << "Spilling " << printReg(UnspilledCSGPR, RegInfo)
2743 << " to get a scratch register.\n");
2744 SavedRegs.set(UnspilledCSGPR);
2745 ExtraCSSpill = UnspilledCSGPR;
2746
2747 // MachO's compact unwind format relies on all registers being stored in
2748 // pairs, so if we need to spill one extra for BigStack, then we need to
2749 // store the pair.
2750 if (producePairRegisters(MF)) {
2751 if (UnspilledCSGPRPaired == AArch64::NoRegister) {
2752 // Failed to make a pair for compact unwind format, revert spilling.
2753 if (produceCompactUnwindFrame(*this, MF)) {
2754 SavedRegs.reset(UnspilledCSGPR);
2755 ExtraCSSpill = AArch64::NoRegister;
2756 }
2757 } else
2758 SavedRegs.set(UnspilledCSGPRPaired);
2759 }
2760 }
2761
2762 // If we didn't find an extra callee-saved register to spill, create
2763 // an emergency spill slot.
2764 if (!ExtraCSSpill || MF.getRegInfo().isPhysRegUsed(ExtraCSSpill)) {
2766 const TargetRegisterClass &RC = AArch64::GPR64RegClass;
2767 unsigned Size = TRI->getSpillSize(RC);
2768 Align Alignment = TRI->getSpillAlign(RC);
2769 int FI = MFI.CreateSpillStackObject(Size, Alignment);
2770 RS->addScavengingFrameIndex(FI);
2771 LLVM_DEBUG(dbgs() << "No available CS registers, allocated fi#" << FI
2772 << " as the emergency spill slot.\n");
2773 }
2774 }
2775
2776 // Adding the size of additional 64bit GPR saves.
2777 CSStackSize += 8 * (SavedRegs.count() - NumSavedRegs);
2778
2779 // A Swift asynchronous context extends the frame record with a pointer
2780 // directly before FP.
2781 if (hasFP(MF) && AFI->hasSwiftAsyncContext())
2782 CSStackSize += 8;
2783
2784 uint64_t AlignedCSStackSize = alignTo(CSStackSize, 16);
2785 LLVM_DEBUG(dbgs() << "Estimated stack frame size: "
2786 << EstimatedStackSize + AlignedCSStackSize << " bytes.\n");
2787
2789 AFI->getCalleeSavedStackSize() == AlignedCSStackSize) &&
2790 "Should not invalidate callee saved info");
2791
2792 // Round up to register pair alignment to avoid additional SP adjustment
2793 // instructions.
2794 AFI->setCalleeSavedStackSize(AlignedCSStackSize);
2795 AFI->setCalleeSaveStackHasFreeSpace(AlignedCSStackSize != CSStackSize);
2796 AFI->setSVECalleeSavedStackSize(ZPRCSStackSize, alignTo(PPRCSStackSize, 16));
2797}
2798
2800 MachineFunction &MF, const TargetRegisterInfo *RegInfo,
2801 std::vector<CalleeSavedInfo> &CSI) const {
2802 bool IsWindows = isTargetWindows(MF);
2803 unsigned StackHazardSize = getStackHazardSize(MF);
2804 // To match the canonical windows frame layout, reverse the list of
2805 // callee saved registers to get them laid out by PrologEpilogInserter
2806 // in the right order. (PrologEpilogInserter allocates stack objects top
2807 // down. Windows canonical prologs store higher numbered registers at
2808 // the top, thus have the CSI array start from the highest registers.)
2809 if (IsWindows)
2810 std::reverse(CSI.begin(), CSI.end());
2811
2812 if (CSI.empty())
2813 return true; // Early exit if no callee saved registers are modified!
2814
2815 // Now that we know which registers need to be saved and restored, allocate
2816 // stack slots for them.
2817 MachineFrameInfo &MFI = MF.getFrameInfo();
2818 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
2819
2820 // Insert VG into the list of CSRs, immediately before LR if saved.
2821 if (requiresSaveVG(MF)) {
2822 CalleeSavedInfo VGInfo(AArch64::VG);
2823 auto It =
2824 find_if(CSI, [](auto &Info) { return Info.getReg() == AArch64::LR; });
2825 if (It != CSI.end())
2826 CSI.insert(It, VGInfo);
2827 else
2828 CSI.push_back(VGInfo);
2829 }
2830
2831 Register LastReg = 0;
2832 int HazardSlotIndex = std::numeric_limits<int>::max();
2833 for (auto &CS : CSI) {
2834 MCRegister Reg = CS.getReg();
2835 const TargetRegisterClass *RC = RegInfo->getMinimalPhysRegClass(Reg);
2836
2837 // Create a hazard slot as we switch between GPR and FPR CSRs.
2839 (!LastReg || !AArch64InstrInfo::isFpOrNEON(LastReg)) &&
2841 assert(HazardSlotIndex == std::numeric_limits<int>::max() &&
2842 "Unexpected register order for hazard slot");
2843 HazardSlotIndex = MFI.CreateStackObject(StackHazardSize, Align(8), true);
2844 LLVM_DEBUG(dbgs() << "Created CSR Hazard at slot " << HazardSlotIndex
2845 << "\n");
2846 AFI->setStackHazardCSRSlotIndex(HazardSlotIndex);
2847 MFI.setIsCalleeSavedObjectIndex(HazardSlotIndex, true);
2848 }
2849
2850 unsigned Size = RegInfo->getSpillSize(*RC);
2851 Align Alignment(RegInfo->getSpillAlign(*RC));
2852 int FrameIdx = MFI.CreateStackObject(Size, Alignment, true);
2853 CS.setFrameIdx(FrameIdx);
2854 MFI.setIsCalleeSavedObjectIndex(FrameIdx, true);
2855
2856 // Grab 8 bytes below FP for the extended asynchronous frame info.
2857 if (hasFP(MF) && AFI->hasSwiftAsyncContext() && Reg == AArch64::FP) {
2858 FrameIdx = MFI.CreateStackObject(8, Alignment, true);
2859 AFI->setSwiftAsyncContextFrameIdx(FrameIdx);
2860 MFI.setIsCalleeSavedObjectIndex(FrameIdx, true);
2861 }
2862 LastReg = Reg;
2863 }
2864
2865 // Add hazard slot in the case where no FPR CSRs are present.
2867 HazardSlotIndex == std::numeric_limits<int>::max()) {
2868 HazardSlotIndex = MFI.CreateStackObject(StackHazardSize, Align(8), true);
2869 LLVM_DEBUG(dbgs() << "Created CSR Hazard at slot " << HazardSlotIndex
2870 << "\n");
2871 AFI->setStackHazardCSRSlotIndex(HazardSlotIndex);
2872 MFI.setIsCalleeSavedObjectIndex(HazardSlotIndex, true);
2873 }
2874
2875 return true;
2876}
2877
2879 const MachineFunction &MF) const {
2881 // If the function has streaming-mode changes, don't scavenge a
2882 // spillslot in the callee-save area, as that might require an
2883 // 'addvl' in the streaming-mode-changing call-sequence when the
2884 // function doesn't use a FP.
2885 if (AFI->hasStreamingModeChanges() && !hasFP(MF))
2886 return false;
2887 // Don't allow register salvaging with hazard slots, in case it moves objects
2888 // into the wrong place.
2889 if (AFI->hasStackHazardSlotIndex())
2890 return false;
2891 return AFI->hasCalleeSaveStackFreeSpace();
2892}
2893
2894/// returns true if there are any SVE callee saves.
2896 int &Min, int &Max) {
2897 Min = std::numeric_limits<int>::max();
2898 Max = std::numeric_limits<int>::min();
2899
2900 if (!MFI.isCalleeSavedInfoValid())
2901 return false;
2902
2903 const std::vector<CalleeSavedInfo> &CSI = MFI.getCalleeSavedInfo();
2904 for (auto &CS : CSI) {
2905 if (AArch64::ZPRRegClass.contains(CS.getReg()) ||
2906 AArch64::PPRRegClass.contains(CS.getReg())) {
2907 assert((Max == std::numeric_limits<int>::min() ||
2908 Max + 1 == CS.getFrameIdx()) &&
2909 "SVE CalleeSaves are not consecutive");
2910 Min = std::min(Min, CS.getFrameIdx());
2911 Max = std::max(Max, CS.getFrameIdx());
2912 }
2913 }
2914 return Min != std::numeric_limits<int>::max();
2915}
2916
2918 AssignObjectOffsets AssignOffsets) {
2919 MachineFrameInfo &MFI = MF.getFrameInfo();
2920 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
2921
2922 SVEStackSizes SVEStack{};
2923
2924 // With SplitSVEObjects we maintain separate stack offsets for predicates
2925 // (PPRs) and SVE vectors (ZPRs). When SplitSVEObjects is disabled predicates
2926 // are included in the SVE vector area.
2927 uint64_t &ZPRStackTop = SVEStack.ZPRStackSize;
2928 uint64_t &PPRStackTop =
2929 AFI->hasSplitSVEObjects() ? SVEStack.PPRStackSize : SVEStack.ZPRStackSize;
2930
2931#ifndef NDEBUG
2932 // First process all fixed stack objects.
2933 for (int I = MFI.getObjectIndexBegin(); I != 0; ++I)
2934 assert(!MFI.hasScalableStackID(I) &&
2935 "SVE vectors should never be passed on the stack by value, only by "
2936 "reference.");
2937#endif
2938
2939 auto AllocateObject = [&](int FI) {
2941 ? ZPRStackTop
2942 : PPRStackTop;
2943
2944 // FIXME: Given that the length of SVE vectors is not necessarily a power of
2945 // two, we'd need to align every object dynamically at runtime if the
2946 // alignment is larger than 16. This is not yet supported.
2947 Align Alignment = MFI.getObjectAlign(FI);
2948 if (Alignment > Align(16))
2950 "Alignment of scalable vectors > 16 bytes is not yet supported");
2951
2952 StackTop += MFI.getObjectSize(FI);
2953 StackTop = alignTo(StackTop, Alignment);
2954
2955 assert(StackTop < (uint64_t)std::numeric_limits<int64_t>::max() &&
2956 "SVE StackTop far too large?!");
2957
2958 int64_t Offset = -int64_t(StackTop);
2959 if (AssignOffsets == AssignObjectOffsets::Yes)
2960 MFI.setObjectOffset(FI, Offset);
2961
2962 LLVM_DEBUG(dbgs() << "alloc FI(" << FI << ") at SP[" << Offset << "]\n");
2963 };
2964
2965 // Then process all callee saved slots.
2966 int MinCSFrameIndex, MaxCSFrameIndex;
2967 if (getSVECalleeSaveSlotRange(MFI, MinCSFrameIndex, MaxCSFrameIndex)) {
2968 for (int FI = MinCSFrameIndex; FI <= MaxCSFrameIndex; ++FI)
2969 AllocateObject(FI);
2970 }
2971
2972 // Ensure the CS area is 16-byte aligned.
2973 PPRStackTop = alignTo(PPRStackTop, Align(16U));
2974 ZPRStackTop = alignTo(ZPRStackTop, Align(16U));
2975
2976 // Create a buffer of SVE objects to allocate and sort it.
2977 SmallVector<int, 8> ObjectsToAllocate;
2978 // If we have a stack protector, and we've previously decided that we have SVE
2979 // objects on the stack and thus need it to go in the SVE stack area, then it
2980 // needs to go first.
2981 int StackProtectorFI = -1;
2982 if (MFI.hasStackProtectorIndex()) {
2983 StackProtectorFI = MFI.getStackProtectorIndex();
2984 if (MFI.getStackID(StackProtectorFI) == TargetStackID::ScalableVector)
2985 ObjectsToAllocate.push_back(StackProtectorFI);
2986 }
2987
2988 for (int FI = 0, E = MFI.getObjectIndexEnd(); FI != E; ++FI) {
2989 if (FI == StackProtectorFI || MFI.isDeadObjectIndex(FI) ||
2991 continue;
2992
2995 continue;
2996
2997 ObjectsToAllocate.push_back(FI);
2998 }
2999
3000 // Allocate all SVE locals and spills
3001 for (unsigned FI : ObjectsToAllocate)
3002 AllocateObject(FI);
3003
3004 PPRStackTop = alignTo(PPRStackTop, Align(16U));
3005 ZPRStackTop = alignTo(ZPRStackTop, Align(16U));
3006
3007 if (AssignOffsets == AssignObjectOffsets::Yes)
3008 AFI->setStackSizeSVE(SVEStack.ZPRStackSize, SVEStack.PPRStackSize);
3009
3010 return SVEStack;
3011}
3012
3014 MachineFunction &MF, RegScavenger *RS) const {
3016 "Upwards growing stack unsupported");
3017
3019
3020 // If this function isn't doing Win64-style C++ EH, we don't need to do
3021 // anything.
3022 if (!MF.hasEHFunclets())
3023 return;
3024
3025 MachineFrameInfo &MFI = MF.getFrameInfo();
3026 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
3027
3028 // Win64 C++ EH needs to allocate space for the catch objects in the fixed
3029 // object area right next to the UnwindHelp object.
3030 WinEHFuncInfo &EHInfo = *MF.getWinEHFuncInfo();
3031 int64_t CurrentOffset =
3033 for (WinEHTryBlockMapEntry &TBME : EHInfo.TryBlockMap) {
3034 for (WinEHHandlerType &H : TBME.HandlerArray) {
3035 int FrameIndex = H.CatchObj.FrameIndex;
3036 if ((FrameIndex != INT_MAX) && MFI.getObjectOffset(FrameIndex) == 0) {
3037 CurrentOffset =
3038 alignTo(CurrentOffset, MFI.getObjectAlign(FrameIndex).value());
3039 CurrentOffset += MFI.getObjectSize(FrameIndex);
3040 MFI.setObjectOffset(FrameIndex, -CurrentOffset);
3041 }
3042 }
3043 }
3044
3045 // Create an UnwindHelp object.
3046 // The UnwindHelp object is allocated at the start of the fixed object area
3047 int64_t UnwindHelpOffset = alignTo(CurrentOffset + 8, Align(16));
3048 assert(UnwindHelpOffset == getFixedObjectSize(MF, AFI, /*IsWin64*/ true,
3049 /*IsFunclet*/ false) &&
3050 "UnwindHelpOffset must be at the start of the fixed object area");
3051 int UnwindHelpFI = MFI.CreateFixedObject(/*Size*/ 8, -UnwindHelpOffset,
3052 /*IsImmutable=*/false);
3053 EHInfo.UnwindHelpFrameIdx = UnwindHelpFI;
3054
3055 MachineBasicBlock &MBB = MF.front();
3056 auto MBBI = MBB.begin();
3057 while (MBBI != MBB.end() && MBBI->getFlag(MachineInstr::FrameSetup))
3058 ++MBBI;
3059
3060 // We need to store -2 into the UnwindHelp object at the start of the
3061 // function.
3062 DebugLoc DL;
3063 RS->enterBasicBlockEnd(MBB);
3064 RS->backward(MBBI);
3065 Register DstReg = RS->FindUnusedReg(&AArch64::GPR64commonRegClass);
3066 assert(DstReg && "There must be a free register after frame setup");
3067 const AArch64InstrInfo &TII =
3068 *MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
3069 BuildMI(MBB, MBBI, DL, TII.get(AArch64::MOVi64imm), DstReg).addImm(-2);
3070 BuildMI(MBB, MBBI, DL, TII.get(AArch64::STURXi))
3071 .addReg(DstReg, getKillRegState(true))
3072 .addFrameIndex(UnwindHelpFI)
3073 .addImm(0);
3074}
3075
3076namespace {
3077struct TagStoreInstr {
3079 int64_t Offset, Size;
3080 explicit TagStoreInstr(MachineInstr *MI, int64_t Offset, int64_t Size)
3081 : MI(MI), Offset(Offset), Size(Size) {}
3082};
3083
3084class TagStoreEdit {
3085 MachineFunction *MF;
3086 MachineBasicBlock *MBB;
3087 MachineRegisterInfo *MRI;
3088 // Tag store instructions that are being replaced.
3090 // Combined memref arguments of the above instructions.
3092
3093 // Replace allocation tags in [FrameReg + FrameRegOffset, FrameReg +
3094 // FrameRegOffset + Size) with the address tag of SP.
3095 Register FrameReg;
3096 StackOffset FrameRegOffset;
3097 int64_t Size;
3098 // If not std::nullopt, move FrameReg to (FrameReg + FrameRegUpdate) at the
3099 // end.
3100 std::optional<int64_t> FrameRegUpdate;
3101 // MIFlags for any FrameReg updating instructions.
3102 unsigned FrameRegUpdateFlags;
3103
3104 // Use zeroing instruction variants.
3105 bool ZeroData;
3106 DebugLoc DL;
3107
3108 void emitUnrolled(MachineBasicBlock::iterator InsertI);
3109 void emitLoop(MachineBasicBlock::iterator InsertI);
3110
3111public:
3112 TagStoreEdit(MachineBasicBlock *MBB, bool ZeroData)
3113 : MBB(MBB), ZeroData(ZeroData) {
3114 MF = MBB->getParent();
3115 MRI = &MF->getRegInfo();
3116 }
3117 // Add an instruction to be replaced. Instructions must be added in the
3118 // ascending order of Offset, and have to be adjacent.
3119 void addInstruction(TagStoreInstr I) {
3120 assert((TagStores.empty() ||
3121 TagStores.back().Offset + TagStores.back().Size == I.Offset) &&
3122 "Non-adjacent tag store instructions.");
3123 TagStores.push_back(I);
3124 }
3125 void clear() { TagStores.clear(); }
3126 // Emit equivalent code at the given location, and erase the current set of
3127 // instructions. May skip if the replacement is not profitable. May invalidate
3128 // the input iterator and replace it with a valid one.
3129 void emitCode(MachineBasicBlock::iterator &InsertI,
3130 const AArch64FrameLowering *TFI, bool TryMergeSPUpdate);
3131};
3132
3133void TagStoreEdit::emitUnrolled(MachineBasicBlock::iterator InsertI) {
3134 const AArch64InstrInfo *TII =
3135 MF->getSubtarget<AArch64Subtarget>().getInstrInfo();
3136
3137 const int64_t kMinOffset = -256 * 16;
3138 const int64_t kMaxOffset = 255 * 16;
3139
3140 Register BaseReg = FrameReg;
3141 int64_t BaseRegOffsetBytes = FrameRegOffset.getFixed();
3142 if (BaseRegOffsetBytes < kMinOffset ||
3143 BaseRegOffsetBytes + (Size - Size % 32) > kMaxOffset ||
3144 // BaseReg can be FP, which is not necessarily aligned to 16-bytes. In
3145 // that case, BaseRegOffsetBytes will not be aligned to 16 bytes, which
3146 // is required for the offset of ST2G.
3147 BaseRegOffsetBytes % 16 != 0) {
3148 Register ScratchReg = MRI->createVirtualRegister(&AArch64::GPR64RegClass);
3149 emitFrameOffset(*MBB, InsertI, DL, ScratchReg, BaseReg,
3150 StackOffset::getFixed(BaseRegOffsetBytes), TII);
3151 BaseReg = ScratchReg;
3152 BaseRegOffsetBytes = 0;
3153 }
3154
3155 MachineInstr *LastI = nullptr;
3156 while (Size) {
3157 int64_t InstrSize = (Size > 16) ? 32 : 16;
3158 unsigned Opcode =
3159 InstrSize == 16
3160 ? (ZeroData ? AArch64::STZGi : AArch64::STGi)
3161 : (ZeroData ? AArch64::STZ2Gi : AArch64::ST2Gi);
3162 assert(BaseRegOffsetBytes % 16 == 0);
3163 MachineInstr *I = BuildMI(*MBB, InsertI, DL, TII->get(Opcode))
3164 .addReg(AArch64::SP)
3165 .addReg(BaseReg)
3166 .addImm(BaseRegOffsetBytes / 16)
3167 .setMemRefs(CombinedMemRefs);
3168 // A store to [BaseReg, #0] should go last for an opportunity to fold the
3169 // final SP adjustment in the epilogue.
3170 if (BaseRegOffsetBytes == 0)
3171 LastI = I;
3172 BaseRegOffsetBytes += InstrSize;
3173 Size -= InstrSize;
3174 }
3175
3176 if (LastI)
3177 MBB->splice(InsertI, MBB, LastI);
3178}
3179
3180void TagStoreEdit::emitLoop(MachineBasicBlock::iterator InsertI) {
3181 const AArch64InstrInfo *TII =
3182 MF->getSubtarget<AArch64Subtarget>().getInstrInfo();
3183
3184 Register BaseReg = FrameRegUpdate
3185 ? FrameReg
3186 : MRI->createVirtualRegister(&AArch64::GPR64RegClass);
3187 Register SizeReg = MRI->createVirtualRegister(&AArch64::GPR64RegClass);
3188
3189 emitFrameOffset(*MBB, InsertI, DL, BaseReg, FrameReg, FrameRegOffset, TII);
3190
3191 int64_t LoopSize = Size;
3192 // If the loop size is not a multiple of 32, split off one 16-byte store at
3193 // the end to fold BaseReg update into.
3194 if (FrameRegUpdate && *FrameRegUpdate)
3195 LoopSize -= LoopSize % 32;
3196 MachineInstr *LoopI = BuildMI(*MBB, InsertI, DL,
3197 TII->get(ZeroData ? AArch64::STZGloop_wback
3198 : AArch64::STGloop_wback))
3199 .addDef(SizeReg)
3200 .addDef(BaseReg)
3201 .addImm(LoopSize)
3202 .addReg(BaseReg)
3203 .setMemRefs(CombinedMemRefs);
3204 if (FrameRegUpdate)
3205 LoopI->setFlags(FrameRegUpdateFlags);
3206
3207 int64_t ExtraBaseRegUpdate =
3208 FrameRegUpdate ? (*FrameRegUpdate - FrameRegOffset.getFixed() - Size) : 0;
3209 LLVM_DEBUG(dbgs() << "TagStoreEdit::emitLoop: LoopSize=" << LoopSize
3210 << ", Size=" << Size
3211 << ", ExtraBaseRegUpdate=" << ExtraBaseRegUpdate
3212 << ", FrameRegUpdate=" << FrameRegUpdate
3213 << ", FrameRegOffset.getFixed()="
3214 << FrameRegOffset.getFixed() << "\n");
3215 if (LoopSize < Size) {
3216 assert(FrameRegUpdate);
3217 assert(Size - LoopSize == 16);
3218 // Tag 16 more bytes at BaseReg and update BaseReg.
3219 int64_t STGOffset = ExtraBaseRegUpdate + 16;
3220 assert(STGOffset % 16 == 0 && STGOffset >= -4096 && STGOffset <= 4080 &&
3221 "STG immediate out of range");
3222 BuildMI(*MBB, InsertI, DL,
3223 TII->get(ZeroData ? AArch64::STZGPostIndex : AArch64::STGPostIndex))
3224 .addDef(BaseReg)
3225 .addReg(BaseReg)
3226 .addReg(BaseReg)
3227 .addImm(STGOffset / 16)
3228 .setMemRefs(CombinedMemRefs)
3229 .setMIFlags(FrameRegUpdateFlags);
3230 } else if (ExtraBaseRegUpdate) {
3231 // Update BaseReg.
3232 int64_t AddSubOffset = std::abs(ExtraBaseRegUpdate);
3233 assert(AddSubOffset <= 4095 && "ADD/SUB immediate out of range");
3234 BuildMI(
3235 *MBB, InsertI, DL,
3236 TII->get(ExtraBaseRegUpdate > 0 ? AArch64::ADDXri : AArch64::SUBXri))
3237 .addDef(BaseReg)
3238 .addReg(BaseReg)
3239 .addImm(AddSubOffset)
3240 .addImm(0)
3241 .setMIFlags(FrameRegUpdateFlags);
3242 }
3243}
3244
3245// Check if *II is a register update that can be merged into STGloop that ends
3246// at (Reg + Size). RemainingOffset is the required adjustment to Reg after the
3247// end of the loop.
3248bool canMergeRegUpdate(MachineBasicBlock::iterator II, unsigned Reg,
3249 int64_t Size, int64_t *TotalOffset) {
3250 MachineInstr &MI = *II;
3251 if ((MI.getOpcode() == AArch64::ADDXri ||
3252 MI.getOpcode() == AArch64::SUBXri) &&
3253 MI.getOperand(0).getReg() == Reg && MI.getOperand(1).getReg() == Reg) {
3254 unsigned Shift = AArch64_AM::getShiftValue(MI.getOperand(3).getImm());
3255 int64_t Offset = MI.getOperand(2).getImm() << Shift;
3256 if (MI.getOpcode() == AArch64::SUBXri)
3257 Offset = -Offset;
3258 int64_t PostOffset = Offset - Size;
3259 // TagStoreEdit::emitLoop might emit either an ADD/SUB after the loop, or
3260 // an STGPostIndex which does the last 16 bytes of tag write. Which one is
3261 // chosen depends on the alignment of the loop size, but the difference
3262 // between the valid ranges for the two instructions is small, so we
3263 // conservatively assume that it could be either case here.
3264 //
3265 // Max offset of STGPostIndex, minus the 16 byte tag write folded into that
3266 // instruction.
3267 const int64_t kMaxOffset = 4080 - 16;
3268 // Max offset of SUBXri.
3269 const int64_t kMinOffset = -4095;
3270 if (PostOffset <= kMaxOffset && PostOffset >= kMinOffset &&
3271 PostOffset % 16 == 0) {
3272 *TotalOffset = Offset;
3273 return true;
3274 }
3275 }
3276 return false;
3277}
3278
3279void mergeMemRefs(const SmallVectorImpl<TagStoreInstr> &TSE,
3281 MemRefs.clear();
3282 for (auto &TS : TSE) {
3283 MachineInstr *MI = TS.MI;
3284 // An instruction without memory operands may access anything. Be
3285 // conservative and return an empty list.
3286 if (MI->memoperands_empty()) {
3287 MemRefs.clear();
3288 return;
3289 }
3290 MemRefs.append(MI->memoperands_begin(), MI->memoperands_end());
3291 }
3292}
3293
3294void TagStoreEdit::emitCode(MachineBasicBlock::iterator &InsertI,
3295 const AArch64FrameLowering *TFI,
3296 bool TryMergeSPUpdate) {
3297 if (TagStores.empty())
3298 return;
3299 TagStoreInstr &FirstTagStore = TagStores[0];
3300 TagStoreInstr &LastTagStore = TagStores[TagStores.size() - 1];
3301 Size = LastTagStore.Offset - FirstTagStore.Offset + LastTagStore.Size;
3302 DL = TagStores[0].MI->getDebugLoc();
3303
3304 Register Reg;
3305 FrameRegOffset = TFI->resolveFrameOffsetReference(
3306 *MF, FirstTagStore.Offset, false /*isFixed*/,
3307 TargetStackID::Default /*StackID*/, Reg,
3308 /*PreferFP=*/false, /*ForSimm=*/true);
3309 FrameReg = Reg;
3310 FrameRegUpdate = std::nullopt;
3311
3312 mergeMemRefs(TagStores, CombinedMemRefs);
3313
3314 LLVM_DEBUG({
3315 dbgs() << "Replacing adjacent STG instructions:\n";
3316 for (const auto &Instr : TagStores) {
3317 dbgs() << " " << *Instr.MI;
3318 }
3319 });
3320
3321 // Size threshold where a loop becomes shorter than a linear sequence of
3322 // tagging instructions.
3323 const int kSetTagLoopThreshold = 176;
3324 if (Size < kSetTagLoopThreshold) {
3325 if (TagStores.size() < 2)
3326 return;
3327 emitUnrolled(InsertI);
3328 } else {
3329 MachineInstr *UpdateInstr = nullptr;
3330 int64_t TotalOffset = 0;
3331 if (TryMergeSPUpdate) {
3332 // See if we can merge base register update into the STGloop.
3333 // This is done in AArch64LoadStoreOptimizer for "normal" stores,
3334 // but STGloop is way too unusual for that, and also it only
3335 // realistically happens in function epilogue. Also, STGloop is expanded
3336 // before that pass.
3337 if (InsertI != MBB->end() &&
3338 canMergeRegUpdate(InsertI, FrameReg, FrameRegOffset.getFixed() + Size,
3339 &TotalOffset)) {
3340 UpdateInstr = &*InsertI++;
3341 LLVM_DEBUG(dbgs() << "Folding SP update into loop:\n "
3342 << *UpdateInstr);
3343 }
3344 }
3345
3346 if (!UpdateInstr && TagStores.size() < 2)
3347 return;
3348
3349 if (UpdateInstr) {
3350 FrameRegUpdate = TotalOffset;
3351 FrameRegUpdateFlags = UpdateInstr->getFlags();
3352 }
3353 emitLoop(InsertI);
3354 if (UpdateInstr)
3355 UpdateInstr->eraseFromParent();
3356 }
3357
3358 for (auto &TS : TagStores)
3359 TS.MI->eraseFromParent();
3360}
3361
3362bool isMergeableStackTaggingInstruction(MachineInstr &MI, int64_t &Offset,
3363 int64_t &Size, bool &ZeroData) {
3364 MachineFunction &MF = *MI.getParent()->getParent();
3365 const MachineFrameInfo &MFI = MF.getFrameInfo();
3366
3367 unsigned Opcode = MI.getOpcode();
3368 ZeroData = (Opcode == AArch64::STZGloop || Opcode == AArch64::STZGi ||
3369 Opcode == AArch64::STZ2Gi);
3370
3371 if (Opcode == AArch64::STGloop || Opcode == AArch64::STZGloop) {
3372 if (!MI.getOperand(0).isDead() || !MI.getOperand(1).isDead())
3373 return false;
3374 if (!MI.getOperand(2).isImm() || !MI.getOperand(3).isFI())
3375 return false;
3376 Offset = MFI.getObjectOffset(MI.getOperand(3).getIndex());
3377 Size = MI.getOperand(2).getImm();
3378 return true;
3379 }
3380
3381 if (Opcode == AArch64::STGi || Opcode == AArch64::STZGi)
3382 Size = 16;
3383 else if (Opcode == AArch64::ST2Gi || Opcode == AArch64::STZ2Gi)
3384 Size = 32;
3385 else
3386 return false;
3387
3388 if (MI.getOperand(0).getReg() != AArch64::SP || !MI.getOperand(1).isFI())
3389 return false;
3390
3391 Offset = MFI.getObjectOffset(MI.getOperand(1).getIndex()) +
3392 16 * MI.getOperand(2).getImm();
3393 return true;
3394}
3395
3396static size_t countAvailableScavengerSlots(LivePhysRegs &LiveRegs,
3398 RegScavenger *RS) {
3399 auto FreeGPRs =
3400 llvm::count_if(AArch64::GPR64RegClass, [&LiveRegs, &MRI](auto Reg) {
3401 return LiveRegs.available(MRI, Reg);
3402 });
3403
3404 size_t NumEmergencySlots = 0;
3405 if (RS)
3406 NumEmergencySlots = RS->getNumScavengingFrameIndices();
3407
3408 return FreeGPRs + NumEmergencySlots;
3409}
3410
3411// Detect a run of memory tagging instructions for adjacent stack frame slots,
3412// and replace them with a shorter instruction sequence:
3413// * replace STG + STG with ST2G
3414// * replace STGloop + STGloop with STGloop
3415// This code needs to run when stack slot offsets are already known, but before
3416// FrameIndex operands in STG instructions are eliminated.
3418 const AArch64FrameLowering *TFI,
3419 RegScavenger *RS) {
3420 bool FirstZeroData;
3421 int64_t Size, Offset;
3422 MachineInstr &MI = *II;
3425 if (&MI == &MBB->instr_back())
3426 return II;
3427 if (!isMergeableStackTaggingInstruction(MI, Offset, Size, FirstZeroData))
3428 return II;
3429
3431 Instrs.emplace_back(&MI, Offset, Size);
3432
3433 constexpr int kScanLimit = 10;
3434 int Count = 0;
3436 NextI != E && Count < kScanLimit; ++NextI) {
3437 MachineInstr &MI = *NextI;
3438 bool ZeroData;
3439 int64_t Size, Offset;
3440 // Collect instructions that update memory tags with a FrameIndex operand
3441 // and (when applicable) constant size, and whose output registers are dead
3442 // (the latter is almost always the case in practice). Since these
3443 // instructions effectively have no inputs or outputs, we are free to skip
3444 // any non-aliasing instructions in between without tracking used registers.
3445 if (isMergeableStackTaggingInstruction(MI, Offset, Size, ZeroData)) {
3446 if (ZeroData != FirstZeroData)
3447 break;
3448 Instrs.emplace_back(&MI, Offset, Size);
3449 continue;
3450 }
3451
3452 // Only count non-transient, non-tagging instructions toward the scan
3453 // limit.
3454 if (!MI.isTransient())
3455 ++Count;
3456
3457 // Just in case, stop before the epilogue code starts.
3458 if (MI.getFlag(MachineInstr::FrameSetup) ||
3460 break;
3461
3462 // Reject anything that may alias the collected instructions.
3463 if (MI.mayLoadOrStore() || MI.hasUnmodeledSideEffects() || MI.isCall())
3464 break;
3465 }
3466
3467 // New code will be inserted after the last tagging instruction we've found.
3468 MachineBasicBlock::iterator InsertI = Instrs.back().MI;
3469
3470 // All the gathered stack tag instructions are merged and placed after
3471 // last tag store in the list. The check should be made if the nzcv
3472 // flag is live at the point where we are trying to insert. Otherwise
3473 // the nzcv flag might get clobbered if any stg loops are present.
3474
3475 // FIXME : This approach of bailing out from merge is conservative in
3476 // some ways like even if stg loops are not present after merge the
3477 // insert list, this liveness check is done (which is not needed).
3479 LiveRegs.addLiveOuts(*MBB);
3480 for (auto I = MBB->rbegin();; ++I) {
3481 MachineInstr &MI = *I;
3482 if (MI == InsertI)
3483 break;
3484 LiveRegs.stepBackward(*I);
3485 }
3486 InsertI++;
3487 if (LiveRegs.contains(AArch64::NZCV))
3488 return InsertI;
3489
3490 // Emitting an MTE loop requires two physical registers (BaseReg and
3491 // SizeReg). If the function is under register pressure, the register
3492 // scavenger will crash trying to allocate them. If we don't have at least
3493 // two free slots (free registers + emergency slots), bail out and fall back
3494 // to the unrolled sequence.
3495 if (countAvailableScavengerSlots(LiveRegs, MBB->getParent()->getRegInfo(),
3496 RS) < 2) {
3497 LLVM_DEBUG(
3498 dbgs() << "Failed to merge MTE stack tagging instructions into loop "
3499 << "due to high register pressure.\n");
3500 return InsertI;
3501 }
3502
3503 llvm::stable_sort(Instrs,
3504 [](const TagStoreInstr &Left, const TagStoreInstr &Right) {
3505 return Left.Offset < Right.Offset;
3506 });
3507
3508 // Make sure that we don't have any overlapping stores.
3509 int64_t CurOffset = Instrs[0].Offset;
3510 for (auto &Instr : Instrs) {
3511 if (CurOffset > Instr.Offset)
3512 return NextI;
3513 CurOffset = Instr.Offset + Instr.Size;
3514 }
3515
3516 // Find contiguous runs of tagged memory and emit shorter instruction
3517 // sequences for them when possible.
3518 TagStoreEdit TSE(MBB, FirstZeroData);
3519 std::optional<int64_t> EndOffset;
3520 for (auto &Instr : Instrs) {
3521 if (EndOffset && *EndOffset != Instr.Offset) {
3522 // Found a gap.
3523 TSE.emitCode(InsertI, TFI, /*TryMergeSPUpdate = */ false);
3524 TSE.clear();
3525 }
3526
3527 TSE.addInstruction(Instr);
3528 EndOffset = Instr.Offset + Instr.Size;
3529 }
3530
3531 const MachineFunction *MF = MBB->getParent();
3532 // Multiple FP/SP updates in a loop cannot be described by CFI instructions.
3533 TSE.emitCode(
3534 InsertI, TFI, /*TryMergeSPUpdate = */
3536
3537 return InsertI;
3538}
3539} // namespace
3540
3542 MachineFunction &MF, RegScavenger *RS = nullptr) const {
3543 for (auto &BB : MF)
3544 for (MachineBasicBlock::iterator II = BB.begin(); II != BB.end();) {
3546 II = tryMergeAdjacentSTG(II, this, RS);
3547 }
3548
3549 // By the time this method is called, most of the prologue/epilogue code is
3550 // already emitted, whether its location was affected by the shrink-wrapping
3551 // optimization or not.
3552 if (!MF.getFunction().hasFnAttribute(Attribute::Naked) &&
3553 shouldSignReturnAddressEverywhere(MF))
3555}
3556
3557/// For Win64 AArch64 EH, the offset to the Unwind object is from the SP
3558/// before the update. This is easily retrieved as it is exactly the offset
3559/// that is set in processFunctionBeforeFrameFinalized.
3561 const MachineFunction &MF, int FI, Register &FrameReg,
3562 bool IgnoreSPUpdates) const {
3563 const MachineFrameInfo &MFI = MF.getFrameInfo();
3564 if (IgnoreSPUpdates) {
3565 LLVM_DEBUG(dbgs() << "Offset from the SP for " << FI << " is "
3566 << MFI.getObjectOffset(FI) << "\n");
3567 FrameReg = AArch64::SP;
3568 return StackOffset::getFixed(MFI.getObjectOffset(FI));
3569 }
3570
3571 // Go to common code if we cannot provide sp + offset.
3572 if (MFI.hasVarSizedObjects() ||
3575 return getFrameIndexReference(MF, FI, FrameReg);
3576
3577 FrameReg = AArch64::SP;
3578 return getStackOffset(MF, MFI.getObjectOffset(FI));
3579}
3580
3581/// The parent frame offset (aka dispFrame) is only used on X86_64 to retrieve
3582/// the parent's frame pointer
3584 const MachineFunction &MF) const {
3585 return 0;
3586}
3587
3588/// Funclets only need to account for space for the callee saved registers,
3589/// as the locals are accounted for in the parent's stack frame.
3591 const MachineFunction &MF) const {
3592 // This is the size of the pushed CSRs.
3593 unsigned CSSize =
3594 MF.getInfo<AArch64FunctionInfo>()->getCalleeSavedStackSize();
3595 // This is the amount of stack a funclet needs to allocate.
3596 return alignTo(CSSize + MF.getFrameInfo().getMaxCallFrameSize(),
3597 getStackAlign());
3598}
3599
3600namespace {
3601struct FrameObject {
3602 bool IsValid = false;
3603 // Index of the object in MFI.
3604 int ObjectIndex = 0;
3605 // Group ID this object belongs to.
3606 int GroupIndex = -1;
3607 // This object should be placed first (closest to SP).
3608 bool ObjectFirst = false;
3609 // This object's group (which always contains the object with
3610 // ObjectFirst==true) should be placed first.
3611 bool GroupFirst = false;
3612
3613 // Used to distinguish between FP and GPR accesses. The values are decided so
3614 // that they sort FPR < Hazard < GPR and they can be or'd together.
3615 unsigned Accesses = 0;
3616 enum { AccessFPR = 1, AccessHazard = 2, AccessGPR = 4 };
3617};
3618
3619class GroupBuilder {
3620 SmallVector<int, 8> CurrentMembers;
3621 int NextGroupIndex = 0;
3622 std::vector<FrameObject> &Objects;
3623
3624public:
3625 GroupBuilder(std::vector<FrameObject> &Objects) : Objects(Objects) {}
3626 void AddMember(int Index) { CurrentMembers.push_back(Index); }
3627 void EndCurrentGroup() {
3628 if (CurrentMembers.size() > 1) {
3629 // Create a new group with the current member list. This might remove them
3630 // from their pre-existing groups. That's OK, dealing with overlapping
3631 // groups is too hard and unlikely to make a difference.
3632 LLVM_DEBUG(dbgs() << "group:");
3633 for (int Index : CurrentMembers) {
3634 Objects[Index].GroupIndex = NextGroupIndex;
3635 LLVM_DEBUG(dbgs() << " " << Index);
3636 }
3637 LLVM_DEBUG(dbgs() << "\n");
3638 NextGroupIndex++;
3639 }
3640 CurrentMembers.clear();
3641 }
3642};
3643
3644bool FrameObjectCompare(const FrameObject &A, const FrameObject &B) {
3645 // Objects at a lower index are closer to FP; objects at a higher index are
3646 // closer to SP.
3647 //
3648 // For consistency in our comparison, all invalid objects are placed
3649 // at the end. This also allows us to stop walking when we hit the
3650 // first invalid item after it's all sorted.
3651 //
3652 // If we want to include a stack hazard region, order FPR accesses < the
3653 // hazard object < GPRs accesses in order to create a separation between the
3654 // two. For the Accesses field 1 = FPR, 2 = Hazard Object, 4 = GPR.
3655 //
3656 // Otherwise the "first" object goes first (closest to SP), followed by the
3657 // members of the "first" group.
3658 //
3659 // The rest are sorted by the group index to keep the groups together.
3660 // Higher numbered groups are more likely to be around longer (i.e. untagged
3661 // in the function epilogue and not at some earlier point). Place them closer
3662 // to SP.
3663 //
3664 // If all else equal, sort by the object index to keep the objects in the
3665 // original order.
3666 return std::make_tuple(!A.IsValid, A.Accesses, A.ObjectFirst, A.GroupFirst,
3667 A.GroupIndex, A.ObjectIndex) <
3668 std::make_tuple(!B.IsValid, B.Accesses, B.ObjectFirst, B.GroupFirst,
3669 B.GroupIndex, B.ObjectIndex);
3670}
3671} // namespace
3672
3674 const MachineFunction &MF, SmallVectorImpl<int> &ObjectsToAllocate) const {
3676
3677 if ((!OrderFrameObjects && !AFI.hasSplitSVEObjects()) ||
3678 ObjectsToAllocate.empty())
3679 return;
3680
3681 const MachineFrameInfo &MFI = MF.getFrameInfo();
3682 std::vector<FrameObject> FrameObjects(MFI.getObjectIndexEnd());
3683 for (auto &Obj : ObjectsToAllocate) {
3684 FrameObjects[Obj].IsValid = true;
3685 FrameObjects[Obj].ObjectIndex = Obj;
3686 }
3687
3688 // Identify FPR vs GPR slots for hazards, and stack slots that are tagged at
3689 // the same time.
3690 GroupBuilder GB(FrameObjects);
3691 for (auto &MBB : MF) {
3692 for (auto &MI : MBB) {
3693 if (MI.isDebugInstr())
3694 continue;
3695
3696 if (AFI.hasStackHazardSlotIndex()) {
3697 std::optional<int> FI = getLdStFrameID(MI, MFI);
3698 if (FI && *FI >= 0 && *FI < (int)FrameObjects.size()) {
3699 if (MFI.getStackID(*FI) == TargetStackID::ScalableVector ||
3701 FrameObjects[*FI].Accesses |= FrameObject::AccessFPR;
3702 else
3703 FrameObjects[*FI].Accesses |= FrameObject::AccessGPR;
3704 }
3705 }
3706
3707 int OpIndex;
3708 switch (MI.getOpcode()) {
3709 case AArch64::STGloop:
3710 case AArch64::STZGloop:
3711 OpIndex = 3;
3712 break;
3713 case AArch64::STGi:
3714 case AArch64::STZGi:
3715 case AArch64::ST2Gi:
3716 case AArch64::STZ2Gi:
3717 OpIndex = 1;
3718 break;
3719 default:
3720 OpIndex = -1;
3721 }
3722
3723 int TaggedFI = -1;
3724 if (OpIndex >= 0) {
3725 const MachineOperand &MO = MI.getOperand(OpIndex);
3726 if (MO.isFI()) {
3727 int FI = MO.getIndex();
3728 if (FI >= 0 && FI < MFI.getObjectIndexEnd() &&
3729 FrameObjects[FI].IsValid)
3730 TaggedFI = FI;
3731 }
3732 }
3733
3734 // If this is a stack tagging instruction for a slot that is not part of a
3735 // group yet, either start a new group or add it to the current one.
3736 if (TaggedFI >= 0)
3737 GB.AddMember(TaggedFI);
3738 else
3739 GB.EndCurrentGroup();
3740 }
3741 // Groups should never span multiple basic blocks.
3742 GB.EndCurrentGroup();
3743 }
3744
3745 if (AFI.hasStackHazardSlotIndex()) {
3746 FrameObjects[AFI.getStackHazardSlotIndex()].Accesses =
3747 FrameObject::AccessHazard;
3748 // If a stack object is unknown or both GPR and FPR, sort it into GPR.
3749 for (auto &Obj : FrameObjects)
3750 if (!Obj.Accesses ||
3751 Obj.Accesses == (FrameObject::AccessGPR | FrameObject::AccessFPR))
3752 Obj.Accesses = FrameObject::AccessGPR;
3753 }
3754
3755 // If the function's tagged base pointer is pinned to a stack slot, we want to
3756 // put that slot first when possible. This will likely place it at SP + 0,
3757 // and save one instruction when generating the base pointer because IRG does
3758 // not allow an immediate offset.
3759 std::optional<int> TBPI = AFI.getTaggedBasePointerIndex();
3760 if (TBPI) {
3761 FrameObjects[*TBPI].ObjectFirst = true;
3762 FrameObjects[*TBPI].GroupFirst = true;
3763 int FirstGroupIndex = FrameObjects[*TBPI].GroupIndex;
3764 if (FirstGroupIndex >= 0)
3765 for (FrameObject &Object : FrameObjects)
3766 if (Object.GroupIndex == FirstGroupIndex)
3767 Object.GroupFirst = true;
3768 }
3769
3770 llvm::stable_sort(FrameObjects, FrameObjectCompare);
3771
3772 int i = 0;
3773 for (auto &Obj : FrameObjects) {
3774 // All invalid items are sorted at the end, so it's safe to stop.
3775 if (!Obj.IsValid)
3776 break;
3777 ObjectsToAllocate[i++] = Obj.ObjectIndex;
3778 }
3779
3780 LLVM_DEBUG({
3781 dbgs() << "Final frame order:\n";
3782 for (auto &Obj : FrameObjects) {
3783 if (!Obj.IsValid)
3784 break;
3785 dbgs() << " " << Obj.ObjectIndex << ": group " << Obj.GroupIndex;
3786 if (Obj.ObjectFirst)
3787 dbgs() << ", first";
3788 if (Obj.GroupFirst)
3789 dbgs() << ", group-first";
3790 dbgs() << "\n";
3791 }
3792 });
3793}
3794
3795/// Emit a loop to decrement SP until it is equal to TargetReg, with probes at
3796/// least every ProbeSize bytes. Returns an iterator of the first instruction
3797/// after the loop. The difference between SP and TargetReg must be an exact
3798/// multiple of ProbeSize.
3800AArch64FrameLowering::inlineStackProbeLoopExactMultiple(
3801 MachineBasicBlock::iterator MBBI, int64_t ProbeSize,
3802 Register TargetReg) const {
3803 MachineBasicBlock &MBB = *MBBI->getParent();
3804 MachineFunction &MF = *MBB.getParent();
3805 const AArch64InstrInfo *TII =
3806 MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
3807 DebugLoc DL = MBB.findDebugLoc(MBBI);
3808
3809 MachineFunction::iterator MBBInsertPoint = std::next(MBB.getIterator());
3810 MachineBasicBlock *LoopMBB = MF.CreateMachineBasicBlock(MBB.getBasicBlock());
3811 MF.insert(MBBInsertPoint, LoopMBB);
3812 MachineBasicBlock *ExitMBB = MF.CreateMachineBasicBlock(MBB.getBasicBlock());
3813 MF.insert(MBBInsertPoint, ExitMBB);
3814
3815 // SUB SP, SP, #ProbeSize (or equivalent if ProbeSize is not encodable
3816 // in SUB).
3817 emitFrameOffset(*LoopMBB, LoopMBB->end(), DL, AArch64::SP, AArch64::SP,
3818 StackOffset::getFixed(-ProbeSize), TII,
3820 // LDR XZR, [SP]
3821 BuildMI(*LoopMBB, LoopMBB->end(), DL, TII->get(AArch64::LDRXui))
3822 .addDef(AArch64::XZR)
3823 .addReg(AArch64::SP)
3824 .addImm(0)
3828 Align(8)))
3830 // CMP SP, TargetReg
3831 BuildMI(*LoopMBB, LoopMBB->end(), DL, TII->get(AArch64::SUBSXrx64),
3832 AArch64::XZR)
3833 .addReg(AArch64::SP)
3834 .addReg(TargetReg)
3837 // B.CC Loop
3838 BuildMI(*LoopMBB, LoopMBB->end(), DL, TII->get(AArch64::Bcc))
3840 .addMBB(LoopMBB)
3842
3843 LoopMBB->addSuccessor(ExitMBB);
3844 LoopMBB->addSuccessor(LoopMBB);
3845 // Synthesize the exit MBB.
3846 ExitMBB->splice(ExitMBB->end(), &MBB, MBBI, MBB.end());
3848 MBB.addSuccessor(LoopMBB);
3849 // Update liveins.
3850 fullyRecomputeLiveIns({ExitMBB, LoopMBB});
3851
3852 return ExitMBB->begin();
3853}
3854
3855void AArch64FrameLowering::inlineStackProbeFixed(
3856 MachineBasicBlock::iterator MBBI, Register ScratchReg, int64_t FrameSize,
3857 StackOffset CFAOffset) const {
3858 MachineBasicBlock *MBB = MBBI->getParent();
3859 MachineFunction &MF = *MBB->getParent();
3860 const AArch64InstrInfo *TII =
3861 MF.getSubtarget<AArch64Subtarget>().getInstrInfo();
3862 AArch64FunctionInfo *AFI = MF.getInfo<AArch64FunctionInfo>();
3863 bool EmitAsyncCFI = AFI->needsAsyncDwarfUnwindInfo(MF);
3864 bool HasFP = hasFP(MF);
3865
3866 DebugLoc DL;
3867 int64_t ProbeSize = MF.getInfo<AArch64FunctionInfo>()->getStackProbeSize();
3868 int64_t NumBlocks = FrameSize / ProbeSize;
3869 int64_t ResidualSize = FrameSize % ProbeSize;
3870
3871 LLVM_DEBUG(dbgs() << "Stack probing: total " << FrameSize << " bytes, "
3872 << NumBlocks << " blocks of " << ProbeSize
3873 << " bytes, plus " << ResidualSize << " bytes\n");
3874
3875 // Decrement SP by NumBlock * ProbeSize bytes, with either unrolled or
3876 // ordinary loop.
3877 if (NumBlocks <= AArch64::StackProbeMaxLoopUnroll) {
3878 for (int i = 0; i < NumBlocks; ++i) {
3879 // SUB SP, SP, #ProbeSize (or equivalent if ProbeSize is not
3880 // encodable in a SUB).
3881 emitFrameOffset(*MBB, MBBI, DL, AArch64::SP, AArch64::SP,
3882 StackOffset::getFixed(-ProbeSize), TII,
3883 MachineInstr::FrameSetup, false, false, nullptr,
3884 EmitAsyncCFI && !HasFP, CFAOffset);
3885 CFAOffset += StackOffset::getFixed(ProbeSize);
3886 // LDR XZR, [SP]
3887 BuildMI(*MBB, MBBI, DL, TII->get(AArch64::LDRXui))
3888 .addDef(AArch64::XZR)
3889 .addReg(AArch64::SP)
3890 .addImm(0)
3894 Align(8)))
3896 }
3897 } else if (NumBlocks != 0) {
3898 // SUB ScratchReg, SP, #FrameSize (or equivalent if FrameSize is not
3899 // encodable in ADD). ScrathReg may temporarily become the CFA register.
3900 emitFrameOffset(*MBB, MBBI, DL, ScratchReg, AArch64::SP,
3901 StackOffset::getFixed(-ProbeSize * NumBlocks), TII,
3902 MachineInstr::FrameSetup, false, false, nullptr,
3903 EmitAsyncCFI && !HasFP, CFAOffset);
3904 CFAOffset += StackOffset::getFixed(ProbeSize * NumBlocks);
3905 MBBI = inlineStackProbeLoopExactMultiple(MBBI, ProbeSize, ScratchReg);
3906 MBB = MBBI->getParent();
3907 if (EmitAsyncCFI && !HasFP) {
3908 // Set the CFA register back to SP.
3909 CFIInstBuilder(*MBB, MBBI, MachineInstr::FrameSetup)
3910 .buildDefCFARegister(AArch64::SP);
3911 }
3912 }
3913
3914 if (ResidualSize != 0) {
3915 // SUB SP, SP, #ResidualSize (or equivalent if ResidualSize is not encodable
3916 // in SUB).
3917 emitFrameOffset(*MBB, MBBI, DL, AArch64::SP, AArch64::SP,
3918 StackOffset::getFixed(-ResidualSize), TII,
3919 MachineInstr::FrameSetup, false, false, nullptr,
3920 EmitAsyncCFI && !HasFP, CFAOffset);
3921 if (ResidualSize > AArch64::StackProbeMaxUnprobedStack) {
3922 // LDR XZR, [SP]
3923 BuildMI(*MBB, MBBI, DL, TII->get(AArch64::LDRXui))
3924 .addDef(AArch64::XZR)
3925 .addReg(AArch64::SP)
3926 .addImm(0)
3930 Align(8)))
3932 }
3933 }
3934}
3935
3936void AArch64FrameLowering::inlineStackProbe(MachineFunction &MF,
3937 MachineBasicBlock &MBB) const {
3938 // Get the instructions that need to be replaced. We emit at most two of
3939 // these. Remember them in order to avoid complications coming from the need
3940 // to traverse the block while potentially creating more blocks.
3941 SmallVector<MachineInstr *, 4> ToReplace;
3942 for (MachineInstr &MI : MBB)
3943 if (MI.getOpcode() == AArch64::PROBED_STACKALLOC ||
3944 MI.getOpcode() == AArch64::PROBED_STACKALLOC_VAR)
3945 ToReplace.push_back(&MI);
3946
3947 for (MachineInstr *MI : ToReplace) {
3948 if (MI->getOpcode() == AArch64::PROBED_STACKALLOC) {
3949 Register ScratchReg = MI->getOperand(0).getReg();
3950 int64_t FrameSize = MI->getOperand(1).getImm();
3951 StackOffset CFAOffset = StackOffset::get(MI->getOperand(2).getImm(),
3952 MI->getOperand(3).getImm());
3953 inlineStackProbeFixed(MI->getIterator(), ScratchReg, FrameSize,
3954 CFAOffset);
3955 } else {
3956 assert(MI->getOpcode() == AArch64::PROBED_STACKALLOC_VAR &&
3957 "Stack probe pseudo-instruction expected");
3958 const AArch64InstrInfo *TII =
3959 MI->getMF()->getSubtarget<AArch64Subtarget>().getInstrInfo();
3960 Register TargetReg = MI->getOperand(0).getReg();
3961 (void)TII->probedStackAlloc(MI->getIterator(), TargetReg, true);
3962 }
3963 MI->eraseFromParent();
3964 }
3965}
3966
3969 NotAccessed = 0, // Stack object not accessed by load/store instructions.
3970 GPR = 1 << 0, // A general purpose register.
3971 PPR = 1 << 1, // A predicate register.
3972 FPR = 1 << 2, // A floating point/Neon/SVE register.
3973 };
3974
3975 int Idx;
3977 int64_t Size;
3978 unsigned AccessTypes;
3979
3981
3982 bool operator<(const StackAccess &Rhs) const {
3983 return std::make_tuple(start(), Idx) <
3984 std::make_tuple(Rhs.start(), Rhs.Idx);
3985 }
3986
3987 bool isCPU() const {
3988 // Predicate register load and store instructions execute on the CPU.
3990 }
3991 bool isSME() const { return AccessTypes & AccessType::FPR; }
3992 bool isMixed() const { return isCPU() && isSME(); }
3993
3994 int64_t start() const { return Offset.getFixed() + Offset.getScalable(); }
3995 int64_t end() const { return start() + Size; }
3996
3997 std::string getTypeString() const {
3998 switch (AccessTypes) {
3999 case AccessType::FPR:
4000 return "FPR";
4001 case AccessType::PPR:
4002 return "PPR";
4003 case AccessType::GPR:
4004 return "GPR";
4006 return "NA";
4007 default:
4008 return "Mixed";
4009 }
4010 }
4011
4012 void print(raw_ostream &OS) const {
4013 OS << getTypeString() << " stack object at [SP"
4014 << (Offset.getFixed() < 0 ? "" : "+") << Offset.getFixed();
4015 if (Offset.getScalable())
4016 OS << (Offset.getScalable() < 0 ? "" : "+") << Offset.getScalable()
4017 << " * vscale";
4018 OS << "]";
4019 }
4020};
4021
4022static inline raw_ostream &operator<<(raw_ostream &OS, const StackAccess &SA) {
4023 SA.print(OS);
4024 return OS;
4025}
4026
4027void AArch64FrameLowering::emitRemarks(
4028 const MachineFunction &MF, MachineOptimizationRemarkEmitter *ORE) const {
4029
4030 auto *AFI = MF.getInfo<AArch64FunctionInfo>();
4032 return;
4033
4034 unsigned StackHazardSize = getStackHazardSize(MF);
4035 const uint64_t HazardSize =
4036 (StackHazardSize) ? StackHazardSize : StackHazardRemarkSize;
4037
4038 if (HazardSize == 0)
4039 return;
4040
4041 const MachineFrameInfo &MFI = MF.getFrameInfo();
4042 // Bail if function has no stack objects.
4043 if (!MFI.hasStackObjects())
4044 return;
4045
4046 std::vector<StackAccess> StackAccesses(MFI.getNumObjects());
4047
4048 size_t NumFPLdSt = 0;
4049 size_t NumNonFPLdSt = 0;
4050
4051 // Collect stack accesses via Load/Store instructions.
4052 for (const MachineBasicBlock &MBB : MF) {
4053 for (const MachineInstr &MI : MBB) {
4054 if (!MI.mayLoadOrStore() || MI.getNumMemOperands() < 1)
4055 continue;
4056 for (MachineMemOperand *MMO : MI.memoperands()) {
4057 std::optional<int> FI = getMMOFrameID(MMO, MFI);
4058 if (FI && !MFI.isDeadObjectIndex(*FI)) {
4059 int FrameIdx = *FI;
4060
4061 size_t ArrIdx = FrameIdx + MFI.getNumFixedObjects();
4062 if (StackAccesses[ArrIdx].AccessTypes == StackAccess::NotAccessed) {
4063 StackAccesses[ArrIdx].Idx = FrameIdx;
4064 StackAccesses[ArrIdx].Offset =
4065 getFrameIndexReferenceFromSP(MF, FrameIdx);
4066 StackAccesses[ArrIdx].Size = MFI.getObjectSize(FrameIdx);
4067 }
4068
4069 unsigned RegTy = StackAccess::AccessType::GPR;
4070 if (MFI.hasScalableStackID(FrameIdx))
4073 RegTy = StackAccess::FPR;
4074
4075 StackAccesses[ArrIdx].AccessTypes |= RegTy;
4076
4077 if (RegTy == StackAccess::FPR)
4078 ++NumFPLdSt;
4079 else
4080 ++NumNonFPLdSt;
4081 }
4082 }
4083 }
4084 }
4085
4086 if (NumFPLdSt == 0 || NumNonFPLdSt == 0)
4087 return;
4088
4089 llvm::sort(StackAccesses);
4090 llvm::erase_if(StackAccesses, [](const StackAccess &S) {
4092 });
4093
4096
4097 if (StackAccesses.front().isMixed())
4098 MixedObjects.push_back(&StackAccesses.front());
4099
4100 for (auto It = StackAccesses.begin(), End = std::prev(StackAccesses.end());
4101 It != End; ++It) {
4102 const auto &First = *It;
4103 const auto &Second = *(It + 1);
4104
4105 if (Second.isMixed())
4106 MixedObjects.push_back(&Second);
4107
4108 if ((First.isSME() && Second.isCPU()) ||
4109 (First.isCPU() && Second.isSME())) {
4110 uint64_t Distance = static_cast<uint64_t>(Second.start() - First.end());
4111 if (Distance < HazardSize)
4112 HazardPairs.emplace_back(&First, &Second);
4113 }
4114 }
4115
4116 auto EmitRemark = [&](llvm::StringRef Str) {
4117 ORE->emit([&]() {
4118 auto R = MachineOptimizationRemarkAnalysis(
4119 "sme", "StackHazard", MF.getFunction().getSubprogram(), &MF.front());
4120 return R << formatv("stack hazard in '{0}': ", MF.getName()).str() << Str;
4121 });
4122 };
4123
4124 for (const auto &P : HazardPairs)
4125 EmitRemark(formatv("{0} is too close to {1}", *P.first, *P.second).str());
4126
4127 for (const auto *Obj : MixedObjects)
4128 EmitRemark(
4129 formatv("{0} accessed by both GP and FP instructions", *Obj).str());
4130}
static void getLiveRegsForEntryMBB(LivePhysRegs &LiveRegs, const MachineBasicBlock &MBB)
static const unsigned DefaultSafeSPDisplacement
This is the biggest offset to the stack pointer we can encode in aarch64 instructions (without using ...
static RegState getPrologueDeath(MachineFunction &MF, unsigned Reg)
static bool produceCompactUnwindFrame(const AArch64FrameLowering &, MachineFunction &MF)
static cl::opt< bool > StackTaggingMergeSetTag("stack-tagging-merge-settag", cl::desc("merge settag instruction in function epilog"), cl::init(true), cl::Hidden)
bool enableMultiVectorSpillFill(const AArch64Subtarget &Subtarget, MachineFunction &MF)
static std::optional< int > getLdStFrameID(const MachineInstr &MI, const MachineFrameInfo &MFI)
static cl::opt< bool > SplitSVEObjects("aarch64-split-sve-objects", cl::desc("Split allocation of ZPR & PPR objects"), cl::init(true), cl::Hidden)
static cl::opt< bool > StackHazardInNonStreaming("aarch64-stack-hazard-in-non-streaming", cl::init(false), cl::Hidden)
void computeCalleeSaveRegisterPairs(const AArch64FrameLowering &AFL, MachineFunction &MF, ArrayRef< CalleeSavedInfo > CSI, const TargetRegisterInfo *TRI, SmallVectorImpl< RegPairInfo > &RegPairs, bool NeedsFrameRecord)
static cl::opt< bool > OrderFrameObjects("aarch64-order-frame-objects", cl::desc("sort stack allocations"), cl::init(true), cl::Hidden)
static cl::opt< bool > DisableMultiVectorSpillFill("aarch64-disable-multivector-spill-fill", cl::desc("Disable use of LD/ST pairs for SME2 or SVE2p1"), cl::init(false), cl::Hidden)
static cl::opt< bool > EnableRedZone("aarch64-redzone", cl::desc("enable use of redzone on AArch64"), cl::init(false), cl::Hidden)
static bool invalidateRegisterPairing(bool SpillExtendedVolatile, unsigned SpillCount, unsigned Reg1, unsigned Reg2, bool UsesWinAAPCS, bool NeedsWinCFI, bool NeedsFrameRecord, const TargetRegisterInfo *TRI)
Returns true if Reg1 and Reg2 cannot be paired using a ldp/stp instruction.
cl::opt< bool > EnableHomogeneousPrologEpilog("homogeneous-prolog-epilog", cl::Hidden, cl::desc("Emit homogeneous prologue and epilogue for the size " "optimization (default = off)"))
static bool isLikelyToHaveSVEStack(const AArch64FrameLowering &AFL, const MachineFunction &MF)
static bool invalidateWindowsRegisterPairing(bool SpillExtendedVolatile, unsigned SpillCount, unsigned Reg1, unsigned Reg2, bool NeedsWinCFI, const TargetRegisterInfo *TRI)
static SVEStackSizes determineSVEStackSizes(MachineFunction &MF, AssignObjectOffsets AssignOffsets)
Process all the SVE stack objects and the SVE stack size and offsets for each object.
static bool isTargetWindows(const MachineFunction &MF)
static unsigned estimateRSStackSizeLimit(MachineFunction &MF)
Look at each instruction that references stack frames and return the stack size limit beyond which so...
static bool getSVECalleeSaveSlotRange(const MachineFrameInfo &MFI, int &Min, int &Max)
returns true if there are any SVE callee saves.
static cl::opt< unsigned > StackHazardRemarkSize("aarch64-stack-hazard-remark-size", cl::init(0), cl::Hidden)
static MCRegister getRegisterOrZero(MCRegister Reg, bool HasSVE)
static unsigned getStackHazardSize(const MachineFunction &MF)
MCRegister findFreePredicateReg(BitVector &SavedRegs)
static bool isPPRAccess(const MachineInstr &MI)
static std::optional< int > getMMOFrameID(MachineMemOperand *MMO, const MachineFrameInfo &MFI)
assert(UImm &&(UImm !=~static_cast< T >(0)) &&"Invalid immediate!")
This file contains the declaration of the AArch64PrologueEmitter and AArch64EpilogueEmitter classes,...
static const int kSetTagLoopThreshold
unsigned Imm
unsigned uint64_t
MachineBasicBlock & MBB
MachineBasicBlock MachineBasicBlock::iterator DebugLoc DL
MachineBasicBlock MachineBasicBlock::iterator MBBI
This file contains the simple types necessary to represent the attributes associated with functions a...
#define CASE(ATTRNAME, AANAME,...)
static GCRegistry::Add< ErlangGC > A("erlang", "erlang-compatible garbage collector")
static GCRegistry::Add< CoreCLRGC > E("coreclr", "CoreCLR-compatible GC")
static GCRegistry::Add< OcamlGC > B("ocaml", "ocaml 3.10-compatible GC")
DXIL Forward Handle Accesses
const HexagonInstrInfo * TII
IRTranslator LLVM IR MI
static std::string getTypeString(Type *T)
Definition LLParser.cpp:68
This file implements the LivePhysRegs utility for tracking liveness of physical registers.
#define F(x, y, z)
Definition MD5.cpp:54
#define I(x, y, z)
Definition MD5.cpp:57
#define H(x, y, z)
Definition MD5.cpp:56
Register Reg
Register const TargetRegisterInfo * TRI
Promote Memory to Register
Definition Mem2Reg.cpp:110
uint64_t IntrinsicInst * II
#define P(N)
This file declares the machine register scavenger class.
static bool contains(SmallPtrSetImpl< ConstantExpr * > &Cache, ConstantExpr *Expr, Constant *C)
Definition Value.cpp:484
This file defines the scope_exit class, which executes user-defined cleanup logic at scope exit.
This file defines the SmallVector class.
#define LLVM_DEBUG(...)
Definition Debug.h:119
StackOffset getSVEStackSize(const MachineFunction &MF) const
Returns the size of the entire SVE stackframe (PPRs + ZPRs).
StackOffset getZPRStackSize(const MachineFunction &MF) const
Returns the size of the entire ZPR stackframe (calleesaves + spills).
void processFunctionBeforeFrameIndicesReplaced(MachineFunction &MF, RegScavenger *RS) const override
processFunctionBeforeFrameIndicesReplaced - This method is called immediately before MO_FrameIndex op...
MachineBasicBlock::iterator eliminateCallFramePseudoInstr(MachineFunction &MF, MachineBasicBlock &MBB, MachineBasicBlock::iterator I) const override
This method is called during prolog/epilog code insertion to eliminate call frame setup and destroy p...
bool canUseAsPrologue(const MachineBasicBlock &MBB) const override
Check whether or not the given MBB can be used as a prologue for the target.
bool enableStackSlotScavenging(const MachineFunction &MF) const override
Returns true if the stack slot holes in the fixed and callee-save stack area should be used when allo...
bool assignCalleeSavedSpillSlots(MachineFunction &MF, const TargetRegisterInfo *TRI, std::vector< CalleeSavedInfo > &CSI) const override
assignCalleeSavedSpillSlots - Allows target to override spill slot assignment logic.
bool spillCalleeSavedRegisters(MachineBasicBlock &MBB, MachineBasicBlock::iterator MI, ArrayRef< CalleeSavedInfo > CSI, const TargetRegisterInfo *TRI) const override
spillCalleeSavedRegisters - Issues instruction(s) to spill all callee saved registers and returns tru...
bool restoreCalleeSavedRegisters(MachineBasicBlock &MBB, MachineBasicBlock::iterator MI, MutableArrayRef< CalleeSavedInfo > CSI, const TargetRegisterInfo *TRI) const override
restoreCalleeSavedRegisters - Issues instruction(s) to restore all callee saved registers and returns...
bool enableFullCFIFixup(const MachineFunction &MF) const override
enableFullCFIFixup - Returns true if we may need to fix the unwind information such that it is accura...
StackOffset getFrameIndexReferenceFromSP(const MachineFunction &MF, int FI) const override
getFrameIndexReferenceFromSP - This method returns the offset from the stack pointer to the slot of t...
bool enableCFIFixup(const MachineFunction &MF) const override
Returns true if we may need to fix the unwind information for the function.
StackOffset getNonLocalFrameIndexReference(const MachineFunction &MF, int FI) const override
getNonLocalFrameIndexReference - This method returns the offset used to reference a frame index locat...
TargetStackID::Value getStackIDForScalableVectors() const override
Returns the StackID that scalable vectors should be associated with.
bool hasFPImpl(const MachineFunction &MF) const override
hasFPImpl - Return true if the specified function should have a dedicated frame pointer register.
void emitPrologue(MachineFunction &MF, MachineBasicBlock &MBB) const override
emitProlog/emitEpilog - These methods insert prolog and epilog code into the function.
void resetCFIToInitialState(MachineBasicBlock &MBB) const override
Emit CFI instructions that recreate the state of the unwind information upon function entry.
bool hasReservedCallFrame(const MachineFunction &MF) const override
hasReservedCallFrame - Under normal circumstances, when a frame pointer is not required,...
bool hasSVECalleeSavesAboveFrameRecord(const MachineFunction &MF) const
StackOffset resolveFrameOffsetReference(const MachineFunction &MF, int64_t ObjectOffset, bool isFixed, TargetStackID::Value StackID, Register &FrameReg, bool PreferFP, bool ForSimm) const
bool canUseRedZone(const MachineFunction &MF) const
Can this function use the red zone for local allocations.
bool needsWinCFI(const MachineFunction &MF) const
bool isFPReserved(const MachineFunction &MF) const
Should the Frame Pointer be reserved for the current function?
void processFunctionBeforeFrameFinalized(MachineFunction &MF, RegScavenger *RS) const override
processFunctionBeforeFrameFinalized - This method is called immediately before the specified function...
int getSEHFrameIndexOffset(const MachineFunction &MF, int FI) const
unsigned getWinEHFuncletFrameSize(const MachineFunction &MF) const
Funclets only need to account for space for the callee saved registers, as the locals are accounted f...
void orderFrameObjects(const MachineFunction &MF, SmallVectorImpl< int > &ObjectsToAllocate) const override
Order the symbols in the local stack frame.
void emitEpilogue(MachineFunction &MF, MachineBasicBlock &MBB) const override
StackOffset getPPRStackSize(const MachineFunction &MF) const
Returns the size of the entire PPR stackframe (calleesaves + spills + hazard padding).
int64_t getArgumentStackToRestore(MachineFunction &MF, MachineBasicBlock &MBB) const
Returns how much of the incoming argument stack area (in bytes) we should clean up in an epilogue.
void determineCalleeSaves(MachineFunction &MF, BitVector &SavedRegs, RegScavenger *RS) const override
This method determines which of the registers reported by TargetRegisterInfo::getCalleeSavedRegs() sh...
StackOffset getFrameIndexReference(const MachineFunction &MF, int FI, Register &FrameReg) const override
getFrameIndexReference - Provide a base+offset reference to an FI slot for debug info.
StackOffset getFrameIndexReferencePreferSP(const MachineFunction &MF, int FI, Register &FrameReg, bool IgnoreSPUpdates) const override
For Win64 AArch64 EH, the offset to the Unwind object is from the SP before the update.
StackOffset resolveFrameIndexReference(const MachineFunction &MF, int FI, Register &FrameReg, bool PreferFP, bool ForSimm) const
unsigned getWinEHParentFrameOffset(const MachineFunction &MF) const override
The parent frame offset (aka dispFrame) is only used on X86_64 to retrieve the parent's frame pointer...
bool requiresSaveVG(const MachineFunction &MF) const
void emitPacRetPlusLeafHardening(MachineFunction &MF) const
Harden the entire function with pac-ret.
AArch64FunctionInfo - This class is derived from MachineFunctionInfo and contains private AArch64-spe...
unsigned getCalleeSavedStackSize(const MachineFrameInfo &MFI) const
void setCalleeSaveBaseToFrameRecordOffset(int Offset)
SignReturnAddress getSignReturnAddressCondition() const
void setStackSizeSVE(uint64_t ZPR, uint64_t PPR)
std::optional< int > getTaggedBasePointerIndex() const
bool needsDwarfUnwindInfo(const MachineFunction &MF) const
void setSVECalleeSavedStackSize(unsigned ZPR, unsigned PPR)
bool needsAsyncDwarfUnwindInfo(const MachineFunction &MF) const
static bool isTailCallReturnInst(const MachineInstr &MI)
Returns true if MI is one of the TCRETURN* instructions.
static bool isFpOrNEON(Register Reg)
Returns whether the physical register is FP or NEON.
const AArch64RegisterInfo * getRegisterInfo() const override
bool isNeonAvailable() const
Returns true if the target has NEON and the function at runtime is known to have NEON enabled (e....
const AArch64InstrInfo * getInstrInfo() const override
const AArch64TargetLowering * getTargetLowering() const override
bool isSVEorStreamingSVEAvailable() const
Returns true if the target has access to either the full range of SVE instructions,...
bool isStreaming() const
Returns true if the function has a streaming body.
bool hasInlineStackProbe(const MachineFunction &MF) const override
True if stack clash protection is enabled for this functions.
unsigned getRedZoneSize(const Function &F) const
Represent a constant reference to an array (0 or more elements consecutively in memory),...
Definition ArrayRef.h:40
size_t size() const
Get the array size.
Definition ArrayRef.h:141
bool empty() const
Check if the array is empty.
Definition ArrayRef.h:136
bool test(unsigned Idx) const
Returns true if bit Idx is set.
Definition BitVector.h:482
BitVector & reset()
Reset all bits in the bitvector.
Definition BitVector.h:409
size_type count() const
Returns the number of bits which are set.
Definition BitVector.h:181
BitVector & set()
Set all bits in the bitvector.
Definition BitVector.h:366
iterator_range< const_set_bits_iterator > set_bits() const
Definition BitVector.h:159
size_type size() const
Returns the number of bits in this bitvector.
Definition BitVector.h:178
Helper class for creating CFI instructions and inserting them into MIR.
The CalleeSavedInfo class tracks the information need to locate where a callee saved register is in t...
A debug info location.
Definition DebugLoc.h:126
bool hasMinSize() const
Optimize this function for minimum size (-Oz).
Definition Function.h:696
CallingConv::ID getCallingConv() const
getCallingConv()/setCallingConv(CC) - These method get and set the calling convention of this functio...
Definition Function.h:273
AttributeList getAttributes() const
Return the attribute list for this Function.
Definition Function.h:329
bool isVarArg() const
isVarArg - Return true if this function takes a variable number of arguments.
Definition Function.h:230
bool hasFnAttribute(Attribute::AttrKind Kind) const
Return true if the function has the attribute.
Definition Function.cpp:730
A set of physical registers with utility functions to track liveness when walking backward/forward th...
bool usesWindowsCFI() const
Definition MCAsmInfo.h:675
Wrapper class representing physical registers. Should be passed by value.
Definition MCRegister.h:41
LLVM_ABI void transferSuccessorsAndUpdatePHIs(MachineBasicBlock *FromMBB)
Transfers all the successors, as in transferSuccessors, and update PHI operands in the successor bloc...
LLVM_ABI iterator getFirstTerminator()
Returns an iterator to the first terminator instruction of this basic block.
LLVM_ABI void addSuccessor(MachineBasicBlock *Succ, BranchProbability Prob=BranchProbability::getUnknown())
Add Succ as a successor of this MachineBasicBlock.
const MachineFunction * getParent() const
Return the MachineFunction containing this basic block.
reverse_iterator rbegin()
iterator insertAfter(iterator I, MachineInstr *MI)
Insert MI into the instruction list after I.
void splice(iterator Where, MachineBasicBlock *Other, iterator From)
Take an instruction from MBB 'Other' at the position From, and insert it into this MBB right before '...
MachineInstrBundleIterator< MachineInstr > iterator
The MachineFrameInfo class represents an abstract stack frame until prolog/epilog code is inserted.
LLVM_ABI int CreateFixedObject(uint64_t Size, int64_t SPOffset, bool IsImmutable, bool isAliased=false)
Create a new object at a fixed location on the stack.
bool hasVarSizedObjects() const
This method may be called any time after instruction selection is complete to determine if the stack ...
const AllocaInst * getObjectAllocation(int ObjectIdx) const
Return the underlying Alloca of the specified stack object if it exists.
LLVM_ABI int CreateStackObject(uint64_t Size, Align Alignment, bool isSpillSlot, const AllocaInst *Alloca=nullptr, uint8_t ID=0)
Create a new statically sized stack object, returning a nonnegative identifier to represent it.
bool hasCalls() const
Return true if the current function has any function calls.
bool isFrameAddressTaken() const
This method may be called any time after instruction selection is complete to determine if there is a...
void setObjectOffset(int ObjectIdx, int64_t SPOffset)
Set the stack frame offset of the specified object.
bool isCalleeSavedObjectIndex(int ObjectIdx) const
uint64_t getMaxCallFrameSize() const
Return the maximum size of a call frame that must be allocated for an outgoing function call.
bool hasPatchPoint() const
This method may be called any time after instruction selection is complete to determine if there is a...
bool hasScalableStackID(int ObjectIdx) const
int getStackProtectorIndex() const
Return the index for the stack protector object.
LLVM_ABI uint64_t estimateStackSize(const MachineFunction &MF) const
Estimate and return the size of the stack frame.
void setStackID(int ObjectIdx, uint8_t ID)
bool isCalleeSavedInfoValid() const
Has the callee saved info been calculated yet?
Align getObjectAlign(int ObjectIdx) const
Return the alignment of the specified stack object.
int64_t getObjectSize(int ObjectIdx) const
Return the size of the specified object.
bool isMaxCallFrameSizeComputed() const
bool hasStackMap() const
This method may be called any time after instruction selection is complete to determine if there is a...
LLVM_ABI int CreateSpillStackObject(uint64_t Size, Align Alignment, TargetStackID::Value StackID=TargetStackID::Default)
Create a new statically sized stack object that represents a spill slot, returning a nonnegative iden...
const std::vector< CalleeSavedInfo > & getCalleeSavedInfo() const
Returns a reference to call saved info vector for the current function.
unsigned getNumObjects() const
Return the number of objects.
int getObjectIndexEnd() const
Return one past the maximum frame object index.
bool hasStackProtectorIndex() const
bool hasStackObjects() const
Return true if there are any stack objects in this function.
uint8_t getStackID(int ObjectIdx) const
unsigned getNumFixedObjects() const
Return the number of fixed objects.
void setIsCalleeSavedObjectIndex(int ObjectIdx, bool IsCalleeSaved)
int64_t getObjectOffset(int ObjectIdx) const
Return the assigned stack offset of the specified object from the incoming stack pointer.
int getObjectIndexBegin() const
Return the minimum frame object index.
void setObjectAlignment(int ObjectIdx, Align Alignment)
setObjectAlignment - Change the alignment of the specified stack object.
bool isDeadObjectIndex(int ObjectIdx) const
Returns true if the specified index corresponds to a dead object.
const WinEHFuncInfo * getWinEHFuncInfo() const
getWinEHFuncInfo - Return information about how the current function uses Windows exception handling.
const TargetSubtargetInfo & getSubtarget() const
getSubtarget - Return the subtarget for which this machine code is being compiled.
LLVM_ABI bool framePointerIsReserved() const
Returns true if the frame pointer must always either point to a new frame record or be un-modified in...
MachineFrameInfo & getFrameInfo()
getFrameInfo - Return the frame info object for the current function.
MachineRegisterInfo & getRegInfo()
getRegInfo - Return information about the registers currently in use.
Function & getFunction()
Return the LLVM function that this machine code represents.
BasicBlockListType::iterator iterator
LLVM_ABI bool disableFramePointerElim() const
Returns true if frame pointer elimination should be disabled for this function.
Ty * getInfo()
getInfo - Keep track of various per-function pieces of information for backends that would like to do...
const MachineBasicBlock & front() const
MachineMemOperand * getMachineMemOperand(MachinePointerInfo PtrInfo, MachineMemOperand::Flags F, LLT MemTy, Align BaseAlignment, const MMOMetadata &Metadata=MMOMetadata(), SyncScope::ID SSID=SyncScope::System, AtomicOrdering Ordering=AtomicOrdering::NotAtomic, AtomicOrdering FailureOrdering=AtomicOrdering::NotAtomic)
getMachineMemOperand - Allocate a new MachineMemOperand.
MachineBasicBlock * CreateMachineBasicBlock(const BasicBlock *BB=nullptr, std::optional< UniqueBBID > BBID=std::nullopt)
CreateMachineInstr - Allocate a new MachineInstr.
void insert(iterator MBBI, MachineBasicBlock *MBB)
const TargetMachine & getTarget() const
getTarget - Return the target machine this machine code is compiled with
const MachineInstrBuilder & setMemRefs(ArrayRef< MachineMemOperand * > MMOs) const
const MachineInstrBuilder & addExternalSymbol(const char *FnName, unsigned TargetFlags=0) const
const MachineInstrBuilder & addReg(Register RegNo, RegState Flags={}, unsigned SubReg=0) const
Add a new virtual register operand.
const MachineInstrBuilder & setMIFlag(MachineInstr::MIFlag Flag) const
const MachineInstrBuilder & addImm(int64_t Val) const
Add a new immediate operand.
const MachineInstrBuilder & addFrameIndex(int Idx) const
const MachineInstrBuilder & addRegMask(const uint32_t *Mask) const
const MachineInstrBuilder & addMBB(MachineBasicBlock *MBB, unsigned TargetFlags=0) const
const MachineInstrBuilder & addDef(Register RegNo, RegState Flags={}, unsigned SubReg=0) const
Add a virtual register definition operand.
const MachineInstrBuilder & setMIFlags(unsigned Flags) const
const MachineInstrBuilder & addMemOperand(MachineMemOperand *MMO) const
Representation of each machine instruction.
void setFlags(unsigned flags)
uint32_t getFlags() const
Return the MI flags bitvector.
LLVM_ABI MachineInstrBundleIterator< MachineInstr > eraseFromParent()
Unlink 'this' from the containing basic block and delete it.
A description of a memory reference used in the backend.
const PseudoSourceValue * getPseudoValue() const
@ MOVolatile
The memory access is volatile.
@ MOLoad
The memory access reads data.
@ MOStore
The memory access writes data.
const Value * getValue() const
Return the base address of the memory access.
MachineOperand class - Representation of each machine instruction operand.
int64_t getImm() const
bool isFI() const
isFI - Tests if this is a MO_FrameIndex operand.
LLVM_ABI void emit(DiagnosticInfoOptimizationBase &OptDiag)
Emit an optimization remark.
MachineRegisterInfo - Keep track of information for virtual and physical registers,...
LLVM_ABI void freezeReservedRegs()
freezeReservedRegs - Called by the register allocator to freeze the set of reserved registers before ...
bool isReserved(MCRegister PhysReg) const
isReserved - Returns true when PhysReg is a reserved register.
LLVM_ABI Register createVirtualRegister(const TargetRegisterClass *RegClass, StringRef Name="")
createVirtualRegister - Create and return a new virtual register in the function with the specified r...
LLVM_ABI bool isLiveIn(Register Reg) const
LLVM_ABI const MCPhysReg * getCalleeSavedRegs() const
Returns list of callee saved registers.
LLVM_ABI bool isPhysRegUsed(MCRegister PhysReg, bool SkipRegMaskTest=false) const
Return true if the specified register is modified or read in this function.
Represent a mutable reference to an array (0 or more elements consecutively in memory),...
Definition ArrayRef.h:294
Wrapper class representing virtual and physical registers.
Definition Register.h:20
constexpr bool isValid() const
Definition Register.h:112
SMEAttrs is a utility class to parse the SME ACLE attributes on functions.
bool hasStreamingInterface() const
bool hasNonStreamingInterfaceAndBody() const
bool hasStreamingBody() const
bool insert(const value_type &X)
Insert a new element into the SetVector.
Definition SetVector.h:157
A SetVector that performs no allocations if smaller than a certain size.
Definition SetVector.h:345
This class consists of common code factored out of the SmallVector class to reduce code duplication b...
reference emplace_back(ArgTypes &&... Args)
void append(ItTy in_start, ItTy in_end)
Add the specified range to the end of the SmallVector.
void push_back(const T &Elt)
This is a 'vector' (really, a variable-sized array), optimized for the case when the array is small.
StackOffset holds a fixed and a scalable offset in bytes.
Definition TypeSize.h:30
int64_t getFixed() const
Returns the fixed component of the stack.
Definition TypeSize.h:46
int64_t getScalable() const
Returns the scalable component of the stack.
Definition TypeSize.h:49
static StackOffset get(int64_t Fixed, int64_t Scalable)
Definition TypeSize.h:41
static StackOffset getScalable(int64_t Scalable)
Definition TypeSize.h:40
static StackOffset getFixed(int64_t Fixed)
Definition TypeSize.h:39
bool hasFP(const MachineFunction &MF) const
hasFP - Return true if the specified function should have a dedicated frame pointer register.
virtual void determineCalleeSaves(MachineFunction &MF, BitVector &SavedRegs, RegScavenger *RS=nullptr) const
This method determines which of the registers reported by TargetRegisterInfo::getCalleeSavedRegs() sh...
int getOffsetOfLocalArea() const
getOffsetOfLocalArea - This method returns the offset of the local area from the stack pointer on ent...
Align getStackAlign() const
getStackAlignment - This method returns the number of bytes to which the stack pointer must be aligne...
StackDirection getStackGrowthDirection() const
getStackGrowthDirection - Return the direction the stack grows
virtual bool enableCFIFixup(const MachineFunction &MF) const
Returns true if we may need to fix the unwind information for the function.
Primary interface to the complete machine description for the target machine.
const Triple & getTargetTriple() const
const MCAsmInfo & getMCAsmInfo() const
Return target specific asm information.
TargetRegisterInfo base class - We assume that the target defines a static array of TargetRegisterDes...
bool hasStackRealignment(const MachineFunction &MF) const
True if stack realignment is required and still possible.
virtual const TargetRegisterInfo * getRegisterInfo() const =0
Return the target's register information.
Triple - Helper class for working with autoconf configuration names.
Definition Triple.h:48
bool isOSBinFormatMachO() const
Tests whether the environment is MachO.
Definition Triple.h:874
This class implements an extremely fast bulk output stream that can only output to a stream.
Definition raw_ostream.h:53
#define llvm_unreachable(msg)
Marks that the current location is not supposed to be reachable.
static unsigned getShiftValue(unsigned Imm)
getShiftValue - Extract the shift value.
static unsigned getArithExtendImm(AArch64_AM::ShiftExtendType ET, unsigned Imm)
getArithExtendImm - Encode the extend type and shift amount for an arithmetic instruction: imm: 3-bit...
const unsigned StackProbeMaxLoopUnroll
Maximum number of iterations to unroll for a constant size probing loop.
const unsigned StackProbeMaxUnprobedStack
Maximum allowed number of unprobed bytes above SP at an ABI boundary.
constexpr char Align[]
Key for Kernel::Arg::Metadata::mAlign.
constexpr char Attrs[]
Key for Kernel::Metadata::mAttrs.
unsigned ID
LLVM IR allows to use arbitrary numbers as calling convention identifiers.
Definition CallingConv.h:24
@ AArch64_SVE_VectorCall
Used between AArch64 SVE functions.
@ PreserveMost
Used for runtime calls that preserves most registers.
Definition CallingConv.h:63
@ CXX_FAST_TLS
Used for access functions.
Definition CallingConv.h:72
@ GHC
Used by the Glasgow Haskell Compiler (GHC).
Definition CallingConv.h:50
@ PreserveAll
Used for runtime calls that preserves (almost) all registers.
Definition CallingConv.h:66
@ Fast
Attempts to make calls as fast as possible (e.g.
Definition CallingConv.h:41
@ PreserveNone
Used for runtime calls that preserves none general registers.
Definition CallingConv.h:90
@ Win64
The C convention as implemented on Windows/x86-64 and AArch64.
@ SwiftTail
This follows the Swift calling convention in how arguments are passed but guarantees tail calls will ...
Definition CallingConv.h:87
@ C
The default llvm calling convention, compatible with C.
Definition CallingConv.h:34
initializer< Ty > init(const Ty &Val)
NodeAddr< InstrNode * > Instr
Definition RDFGraph.h:389
BaseReg
Stack frame base register. Bit 0 of FREInfo.Info.
Definition SFrame.h:77
This is an optimization pass for GlobalISel generic memory operations.
@ Offset
Definition DWP.cpp:577
void stable_sort(R &&Range)
Definition STLExtras.h:2116
MachineInstrBuilder BuildMI(MachineFunction &MF, const MIMetadata &MIMD, const MCInstrDesc &MCID)
Builder interface. Specify how to create the initial instruction itself.
int isAArch64FrameOffsetLegal(const MachineInstr &MI, StackOffset &Offset, bool *OutUseUnscaledOp=nullptr, unsigned *OutUnscaledOp=nullptr, int64_t *EmittableOffset=nullptr)
Check if the Offset is a valid frame offset for MI.
@ Unknown
Not known to have no common set bits.
RegState
Flags to represent properties of register accesses.
@ Define
Register definition.
constexpr RegState getKillRegState(bool B)
decltype(auto) dyn_cast(const From &Val)
dyn_cast<X> - Return the argument parameter cast to the specified type.
Definition Casting.h:643
@ AArch64FrameOffsetCannotUpdate
Offset cannot apply.
constexpr T alignDown(U Value, V Align, W Skew=0)
Returns the largest unsigned integer less than or equal to Value and is Skew mod Align.
Definition MathExtras.h:541
auto dyn_cast_or_null(const Y &Val)
Definition Casting.h:753
bool any_of(R &&range, UnaryPredicate P)
Provide wrappers to std::any_of which take ranges instead of having to pass begin/end explicitly.
Definition STLExtras.h:1746
auto formatv(bool Validate, const char *Fmt, Ts &&...Vals)
auto reverse(ContainerTy &&C)
Definition STLExtras.h:407
void sort(IteratorTy Start, IteratorTy End)
Definition STLExtras.h:1636
LLVM_ABI raw_ostream & dbgs()
dbgs() - This returns a reference to a raw_ostream for debugging messages.
Definition Debug.cpp:209
void emitFrameOffset(MachineBasicBlock &MBB, MachineBasicBlock::iterator MBBI, const DebugLoc &DL, unsigned DestReg, unsigned SrcReg, StackOffset Offset, const TargetInstrInfo *TII, MachineInstr::MIFlag=MachineInstr::NoFlags, bool SetNZCV=false, bool NeedsWinCFI=false, bool *HasWinCFI=nullptr, bool EmitCFAOffset=false, StackOffset InitialOffset={}, unsigned FrameReg=AArch64::SP)
emitFrameOffset - Emit instructions as needed to set DestReg to SrcReg plus Offset.
LLVM_ABI void report_fatal_error(Error Err, bool gen_crash_diag=true)
Definition Error.cpp:163
constexpr uint64_t alignTo(uint64_t Size, Align A)
Returns a multiple of A needed to store Size bytes.
Definition Alignment.h:144
constexpr RegState getDefRegState(bool B)
class LLVM_GSL_OWNER SmallVector
Forward declaration of SmallVector so that calculateSmallVectorDefaultInlinedElements can reference s...
@ First
Helpers to iterate all locations in the MemoryEffectsBase class.
Definition ModRef.h:74
uint16_t MCPhysReg
An unsigned integer type large enough to represent all physical registers, but not necessarily virtua...
Definition MCRegister.h:21
RelativeUniformCounterPtr ValuesPtrExpr VTableAddr Count
Definition InstrProf.h:145
raw_ostream & operator<<(raw_ostream &OS, const APFixedPoint &FX)
auto count_if(R &&Range, UnaryPredicate P)
Wrapper function around std::count_if to count the number of times an element satisfying a given pred...
Definition STLExtras.h:2019
auto find_if(R &&Range, UnaryPredicate P)
Provide wrappers to std::find_if which take ranges instead of having to pass begin/end explicitly.
Definition STLExtras.h:1772
void erase_if(Container &C, UnaryPredicate P)
Provide a container algorithm similar to C++ Library Fundamentals v2's erase_if which is equivalent t...
Definition STLExtras.h:2192
bool is_contained(R &&Range, const E &Element)
Returns true if Element is found in Range.
Definition STLExtras.h:1947
LLVM_ABI const Value * getUnderlyingObject(const Value *V, unsigned MaxLookup=MaxLookupSearchDepth)
This method strips off any GEP address adjustments, pointer casts or llvm.threadlocal....
void fullyRecomputeLiveIns(ArrayRef< MachineBasicBlock * > MBBs)
Convenience function for recomputing live-in's for a set of MBBs until the computation converges.
LLVM_ABI Printable printReg(Register Reg, const TargetRegisterInfo *TRI=nullptr, unsigned SubIdx=0, const MachineRegisterInfo *MRI=nullptr)
Prints virtual and physical registers with or without a TRI instance.
MCRegisterClass TargetRegisterClass
Definition FastISel.h:58
void swap(llvm::BitVector &LHS, llvm::BitVector &RHS)
Implement std::swap in terms of BitVector swap.
Definition BitVector.h:880
bool operator<(const StackAccess &Rhs) const
void print(raw_ostream &OS) const
int64_t start() const
std::string getTypeString() const
int64_t end() const
This struct is a compact representation of a valid (non-zero power of two) alignment.
Definition Alignment.h:39
constexpr uint64_t value() const
This is a hole in the type system and should not be abused.
Definition Alignment.h:77
Pair of physical register and lane mask.
static LLVM_ABI MachinePointerInfo getUnknownStack(MachineFunction &MF)
Stack memory without other information.
static LLVM_ABI MachinePointerInfo getFixedStack(MachineFunction &MF, int FI, int64_t Offset=0)
Return a MachinePointerInfo record that refers to the specified FrameIndex.
SmallVector< WinEHTryBlockMapEntry, 4 > TryBlockMap
SmallVector< WinEHHandlerType, 1 > HandlerArray