| Commit message (Collapse) | Author | Age | Files | Lines |
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Some segments, like the 'vectors' one, will reference code that is
outside of its mapping. But in some other configurations, segments
cannot make these cross-mapping references so happily. Imagine:
.segment "SWAPPABLE"
.proc foo
rts
.endproc
.segment "FIXED"
jsr foo
Here the assembler will properly detect the address of 'foo' in the
context of the 'SWAPPABLE' segment. But what this assembler doesn't know
is that this segment is swappable (e.g. UNROM chip). Hence, if the bank
being mapped right now is not the one containing the 'SWAPPABLE'
segment, then the address computed for 'foo' and used in that 'jsr'
instruction will point to something else entirely. This would be similar
to a use-after-free bug.
This is something that can only be inspected at runtime, and so the
assembler cannot be of much help here. Hence, this commit adds a warning
so the programmer can understand the potentially dangerous operation.
All of that being said, this commit also adds support for "asan:safe" or
"check:safe", which is a magic comment that the programmer can write to
re-assure the assembler that this operation is fine (e.g. there is a
guarantee that the mapped bank is that one we are expecting). Hence, the
code above could now be written like so:
jsr foo ; check:safe
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
| |
And also adjust the code so the clippy from the 2024 edition is fine
with it.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
| |
The enum type for macro arguments has been adjusted so it accepts a
String. This is then used to store the name of the macro for this
argument.
This is a bit of a circlejerk, but in the end it stems from the fact
that macro arguments were sort of a hack defined ad-hoc each time they
were needed. That being said, if we want to report macro arguments being
unused, we need to reference the actual macro definition, not the macro
call.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
| |
It was downright lazy from my end to not inspect the line of the object
that was being unused. Fix this by trying to fetch the bundle's node.
This right now applies to variable definitions. Objects like macro
arguments are pending to be done.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
In the following code:
.macro MACRO ADDR
adc ADDR
.endmacro
MACRO label
label:
nop
The evaluation of the 'adc' instruction went into being re-sized to 1
despite indications of being an address (and hence hinting to a size of
2). This was done because of the optimization to shrink a single byte
left arm from absolute to relative addressing. But on unresolved bundles
this is bad, as everything may still be zeroed out, and the instruction
itself doesn't hint on the end size. This shrinking would then mess with
the segment's offset, and we would end up with an offset of 2 instead of
3 for the label 'label'.
Note that this is in the same spirit as commit ed3d34b0afb1 ("Only
shrink the addressing on resolved bundles"), but now applied to the
get_from_left() function which apparently was spared for whatever
reason.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
When warning users on unused objects (e.g. variables, proc's), we have
to pick up the source of origin for this warning. In places where the
original PNode is not available, then we go with the last source we have
at hand, as that's the usual way to go (i.e. the error occurred at the
current source/context).
That being said, for unused objects that's not desirable, because we
might otherwise claim the error to happen on the file we first
targetted, but it's way more useful to understand where the object was
defined.
Hence, when defining a variable, address or proc, let's actually store
the PNode associated with it as well. This way the checker can hopefully
get the source from this PNode and be more clear to the user.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
| |
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
| |
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
| |
Since the addition of the check() function in commit 6a9dea943c07 ("Add
a new 'check' stage in assemble()"), it can be annoying to some users
that the check for unused/unreferenced object applies by default. Hence,
add a flag so this can be disabled.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
| |
This is actually an expansion from asan(), but it also checks for stuff
that goes outside of the realm of address sanitation. In this aspect,
asan() has been integrated inside of the new check() function.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
We cannot safely detect all scenarios in which there is dead code, but
we can safely do it for .proc's. Even if they are not called directly,
they might be referenced via jump tables and shenanigans like that. Long
story short, if you are not referencing a .proc in any meaningful way,
then we have dead code.
A pattern could also be given in which a function that was too long has
been splitted into smaller functions which are not called
directly. This is, in my opinion, an anti-pattern (as we generally
expect an rts or a jmp from .proc's), and anyways can be avoided via the
use of __fallthrough__ if the programmer is really set to write this
kind of code.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
| |
This allows the programmer to tell nasm to leave each segment into its
own .out file. This can be useful when writing some specific segments
into different chips.
The caveat is that filling is not applied, and hence it's up to the
programmer to account for that if two or more segments happened to be on
the same mapping and needed some alignment in between.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
| |
This is applied whenever a given memory range name is not under the
global context. The end result is that scopes will be shown into the
memory.txt file when using '--write-info'.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
| |
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
| |
Fixes 4f1a9c660108 ("Add the .fallthrough control statement")
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
| |
By default everything is little endian, and that should always be the
case. That being said, if a literal is better expressed in big endian
format for whatever reason, you can now use the .be or the .bigendian
control statements. The terribly named .dbyt control statement can also
be used for this purpose, but that's just to be compatible with the
implementation from ca65.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
| |
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
| |
This mostly involves applying the collapsible_match rule[1] where
appropiate, and using sort_by_key() as pointed out by the latest stable
version of clippy.
[1] https://rust-lang.github.io/rust-clippy/rust-1.95.0/index.html#collapsible_match
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
| |
Since commit 31dd071e927a ("Change .fallthrough to __fallthrough__")
this pseudo control statement is declared via underscores, not like a
proper control statement. That was done in order to avoid clashes with
ca65, and in order to allow nasm's --prelude to produce a .macro
replacement for it when using ca65.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Some macro calls could contain addresses, like this:
MACRO_CALL @address
Before this commit this was resolved at the "crunch" stage, but only
because we were grabbing the value from the part that was not
resolved. However, if this unresolved object came from a node which is
not just a value (e.g. an arithmetic operation), then it just applied
the computation from before the "crunch" stage. So, for something like
this:
MACRO_CALL @address + 1
the argument would have been computed as simply "1", as that unresolved
object only contained the "+ 1" known part.
Fix this by re-evaluating the node at "crunch" stage whenever we have an
attached "node" member to the given pending definition. This way, we can
evaluate the full node and get the proper value from the arithmetic
computation.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Imagine the following scenario:
lda #.lobyte(@address + 1)
Here we have an immediate which comes from the .lobyte call. That being
said, since this is using the '#' symbol, we look ahead after noticing
it just in case we need to disambiguate some scenarios. When a
parenthesis is found, in extract_paren_expression() we temporarily
discard the looking ahead status as expressions inside of parenthesis
are no longer affected by any hypothetical external ambiguity. This,
though, was not being done in parse_arguments(), which is the function
that would be called in the above scenario.
Hence, ensure that we temporarily decrease the look ahead status for the
inner arguments, and then increase it back when we are done. This way,
when the parser finds '@address', it will no longer go the "we have
found a 'value', refrain from ambiguities and return early" route, and
it will check for any binary expression.
Note that all this dance actually applied here, as binary operations are
a source of ambiguities, and the look ahead mechanism was brought about
because of these kinds of expressions. Hence, it's not like the look
ahead mechanism is troublesome in on itself, we just missed the special
handling on this specific scenario.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
| |
This statement allows programmers to reserve a memory range for the
stack. This is useful to reserve a memory region which is not attached
to any variable, and:
1. It is a way to safely reserve space on the $0100 page.
2. The --stats/--write-info flags will be able to add up the memory
space reserved for the stack.
3. Future tooling might be able to use this value as a way to detect
stack overflows.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Bundle call arguments were fine most of the times, when arguments could
be processed as-is and there were no issues with arguments being
overwritten by successive calls.
This was not the case, though, whenever a given instruction was delayed
into a PendingNode status. In this case, the argument would get the
last value, and in some extreme cases that definition might not have
been there any more.
Prevent all of this by providing a list of PendingDefine's, which are a
way to re-create these call-only arguments as they were initially
found. These PendingDefine's are then created on "crunch" on each
PendingNode.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
| |
These are usually capitalized, and thus they are not going to be
following the usual zp_/m_/wr_ pattern.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
| |
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
| |
They should be expanded to a 16-bit address from the zero-page. The fact
that the assembler was complaining about it was simply erroneous.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
| |
There were two bugs involved at the same time. First, for some reason,
the "matches" clause was negated, which defeats the purpose of the
check. This even resulted in a bad test run which was accepted because I
just did not caught it before.
Second, pure indirect addressing mode is only available for the "jmp"
instruction. Hence, there's no need for that "matches" at all, and even
less to filter that based on "jsr" which doesn't even implement this
addressing mode.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
| |
This was already the case for plain addresses, but it was not being
considered in the case of an arithmetic operation.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
In instructions like:
lda address + 1, x
The assembler was picking the size from the "+ 1" side, making this part
of a one-byte size and, hence, assuming it was zeropage indirect
addressing instead of an absolute one. This meant that one byte was cut
off the end binary, and the address was wrong altogether.
Fix this by telling the assembler to assume a .size = 2; for variables
which have been defined as addresses. Then, the operation evaluation can
take that as an indication to set the size to 2 as well and, hence,
assume absolute indirect addressing.
Moreover, this bug was intertwined with another one, where arithmetic
could be borked when both operands were of an unexpected size. For
example:
lda #(300 - 260)
Here the size would've been resolved to 2 as well, but it would have
also reported an error because both operands were 2-bytes long, despite
the end result fitting on a single byte. Since this is an explicit
immediate addressing, then it would have assumed the programmer to try
that with a 16-bit value, which is wrong. This commit also fixes this
bug as it was deeply intertwined to how numbers are computed on
operations and how sizes are evaluated.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
| |
Add the target and effective addresses into the message, so it's more
clear how different they are.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
| |
This was apparently neglected and you were able to pick invalid
identifiers to identify procs, macros and scopes. Ensure this does not
happen again and provide tests for it.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
| |
Implementing it as a control statement has the bad thing that it's
impossible to be compatible with other assemblers such as ca65, as you
cannot create dummy control statements or something like that in
there. Instead of that, we define it with the special underscores which
are still valid for identifiers, and give a "compiler-specific thingie"
flair to it.
This also has the benefit that the parser can be a bit more strict.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
This is exclusive to 'nasm' and it allows the developer to explicitly
tell the assembler than a "fall through" condition is actually desired:
it's not a mistake.
This comes in two flavors. The first, without arguments, just makes this
explicit without much enforcement. The second allows you to pass an
argument which is the name of the function or label you are expecting to
fall through. The assembler will error out if the fall through address
is not the expected one, hence telling the programmer whenever the fall
through condition they thought in the past is no longer true (e.g. the
function has moved somewhere else in the code).
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
Sometimes performing some 'jmp'/'jsr' can be quite pointless, and the
programmer might not be fully aware of this because of the layout of the
code. Imagine:
.proc foo
;; code
jmp bar
.endproc
;; Documentation, comments, extra space, etc.
.proc bar
;; whatever
.endproc
The 'jmp' in the code above tries to perform a call stack optimization,
but it's actually not needed because the next instruction after 'jmp' is
the one inside of 'bar', but that's obfuscated because of the layout.
In these sort of cases (and also for 'jsr' and branches) warn the
programmer about it so it can remove that instruction.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
| |
The way macro expansions work is that the evaluated bundles replace the
current node, but this throws away the context of the line, the source,
etc. from the original line of code.
Add a stack that is pushed/popped when expanding macros, and are then
cloned for each pending node and error. This way, errors have full
context of the original code and can display backtraces for macro
expansion.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
| |
There was a condition race in which some conflicts would be caught but
sometimes wouldn't depending on how/when `memory.memory_ranges` was
being filled.
Instead of this, just fill this vector with otherwise valid
candidates (other checks like "is it used?" still apply in this
context), and then perform the conflict check upon the already filled
vector.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
This way, if you define a constant like:
MY_BUFFER_LEN_IN_BYTES = $10
You can then declare your buffer like so:
zp_buffer = $00 ; asan:reserve MY_BUFFER_LEN_IN_BYTES
And then further in the code you can rely on just using the constant for
bound checking, and then the address sanitizer will check on bound
checks via static analysis as well.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
| |
These addresses are merely references to real APU/PPU addresses which
can be used for ease of use.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
|
| |
When evaluating an arithmetic operation, check on whether it
theoretically spans more than one byte, and set it on the resulting
Bundle.size. It's up to the caller to decide on whether that makes sense
or not. This fixes situations in which:
sta $200 + 1
was marked as having 2 bytes for the size as only the "1" was being
considered for the size on this operation. This would then signal the
"sta" evaluation that it was zeropage instead of absolute addressing.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
| |
This allows for reserving memory regions which go outside of the page
boundary.
Signed-off-by: Miquel Sabaté Solà <mssola@mssola.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
|
| |
Negative variables make no sense when it comes to do stuff like:
variable = -1
lda variable
That being said, as an assembler you never know the hacks and nonsense
programmers are willing to endure. But we do know that if the address
sanitizer is enabled, since if it's an "asan-friendly", then an explicit
negative variable can be barred.
Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
|
| |
|
|
|
|
| |
The check was not being applied on certain conditions.
Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
|
| |
|
|
|
|
|
| |
This way instructions which might make use of bare memory numbers can
freely ignore the address sanitizer when it actually makes sense.
Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
|
| |
|
|
|
|
|
| |
This check ensures that asan-friendly names actually match their
expected scope.
Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
|
| |
|
|
|
|
|
|
|
|
|
|
|
| |
The address sanitizer is now able to detect whenever in an instruction a
memory access is done without using variables. This is now detected for
all instructions except for branching, which falls outside of this
scope.
Moreover, simple arithmetics is allowed and bounds are checked for
simple cases. That being said, more involved bound checks should be done
with other tools (e.g. emulators).
Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
|
| |
|
|
| |
Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
|
| |
|
|
| |
Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
|
| |
|
|
|
|
|
|
| |
This is the initial support for both directives for the address
sanitizer. Note that asan:weak has been moved into asan:ignore, which is
not exactly the same but for now it should suffice.
Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
|
| |
|
|
|
|
|
|
|
|
| |
This commit introduces the ability to inspect the temptative header
before producing the actual output, and with that it checks whether the
Working RAM is being advertised or not. If it is not being advertised
but the assembler detected memory accesses to that region, then we are
in trouble and we should error out.
Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
|
| |
|
|
|
|
|
| |
This means removing a lot of `pub` structs or enums, as well as adding
documentation on `pub` structs.
Signed-off-by: Miquel Sabaté Solà <mikisabate@gmail.com>
|